Blog
Arguments, mostly.
Agent governance, AI regulation, evaluation and the engineering underneath. Written to be disagreed with rather than to rank for anything.
By topic
Everything, newest first
-
The Check Passes. We Still Do Not Believe It.
A National Insurance number has no check digit. HMRC has never published the one for a UTR. Building the UK detection profile meant deciding what a scanner m...
-
A Regulator Wrote Down What an Agent Is
DIFC Regulation 10 governs personal data processed through autonomous systems. It describes an agent as substantially similar to an employee, and requires it...
-
RotaGrant Is Live
The agent governance platform we have been writing about is running, and you can sign in to it. What it does, what it refuses to claim, and how to see it on ...
-
The Agent That Did the Right Thing Ten Thousand Times
In grid operations the dangerous agent is not the one that gets it wrong. It is the one that gets it right at a scale nobody bounded.
-
Nobody Is Making You Govern Your Agents. Do It Anyway.
Agent governance is sold to regulated enterprises. The companies shipping the most agents, into the most accounts, have no regulator, and three better reasons.
-
What an Incident Review Actually Asks For
Not your dashboards. Four questions, in a fixed order, and the third one is the one nobody can answer about an autonomous system.
-
A Service Account Is Not a Person
ALCOA has held for thirty years and every part of it assumes a human actor. What attributable means when the thing writing to your quality system decides for...
-
Delegation Multiplies Authority Unless Something Stops It
Every agent in a multi-agent pipeline is inside its own limit. Nobody is tracking what the tree can spend, and the tree is what has your money.
-
The Claim That Paid Twice
A claims pipeline where every limit was correct and the total was not. What insurers should ask their vendors about delegated authority before the pipeline g...
-
Your Agent Does Not Need a Bigger Context Window
Every document an agent retrieves is an instruction it was never told to distrust. The fix is not more context. It is knowing which parts of it someone else ...
-
The August Deadline Is a Documentation Deadline
High-risk obligations under the EU AI Act land in August. Almost nothing they ask for is a model property. Most of it asks who authorised what, and when.
-
Structured Output Isn't Reliable Output
JSON mode, function calling and constrained decoding give you schema compliance, not semantic reliability. Valid JSON can be completely wrong.
-
The Insurance Industry's AI Blind Spot: Claims Automation Without Trust Infrastructure
Insurance companies are racing to automate claims with AI. Nobody has built for the regulator, the litigant, or the appeals board. That is the blind spot.
-
What Moltbook Reveals About Multi-Agent Trust at Scale
Moltbook isn't an enterprise product - but the vulnerabilities it exposes matter for any organization deploying multi-agent AI systems.
-
The Agent Watchtower, Part 5: Reference Architecture
A complete, implementable design for enterprise agent governance. Concrete specifications, integration patterns, and implementation roadmap.
-
The Agent Watchtower, Part 4: Economics of Agent Operations
The financial model for sustainable AI governance. Cost cascading, ROI-driven routing, and why governance pays for itself.
-
The Agent Watchtower, Part 3: The Autonomy Spectrum
How to balance business unit freedom with enterprise governance. Federated control, trust-based permissions, and why guardrails beat gates.
-
The Agent Watchtower, Part 2: Anatomy of an Agent Control Plane
The technical architecture for unified agent governance: registry, observability, policy and control, and how they make multi-cloud governance possible.
-
The Agent Watchtower, Part 1: The Fragmentation Tax
Banks are deploying AI agents across AWS, Azure, GCP and open-source frameworks. The result: governance blind spots and a ticking regulatory problem.
-
The Eval Crisis: Why Most Benchmarks Don't Matter
Your model scores 90% on MMLU. It still fails in production. The benchmarks everyone obsesses over measure the wrong things for enterprise AI.
-
5 Evals Every Production LLM Needs
Forget MMLU scores. These are the evaluations that actually predict whether your LLM will work in production.
-
The Real Reason Your RAG App Hallucinates (It's Not Chunking)
Everyone's optimizing chunk size and embedding models. The problem is upstream. Your data pipeline strips context before it ever reaches the vector store.
-
Your AI Architecture is Bleeding Money
Cost-per-token is the wrong metric. The real savings come from architectural decisions most teams get wrong.
-
The AI Production Readiness Checklist
The comprehensive checklist for launching LLM-powered features. Evaluation, monitoring, fallbacks, cost controls, and incident response.
-
Prompt Injection is an Unsolved Problem (Here's How to Mitigate Anyway)
There's no complete solution to prompt injection. Here's the defense-in-depth playbook for production AI systems.
-
When to Use Agents vs Deterministic Workflows: A Decision Framework
A concrete decision tree for when to reach for AI agents vs traditional orchestration. Cost, latency, reliability, and compliance dimensions.
-
From 11% to 88% GPU Utilization: How We Built 8x Faster LLM Inference
PyTorch leaves 89% of GPU bandwidth on the table. We fixed it with custom Triton kernels. Here's what we learned building Accelerate.
-
Eval Debt Will End Careers
Tech debt is slow. Eval debt is sudden. The teams that survive will treat evals like unit tests: written first, run always.
-
Agentic AI Is a Cost Center, Not a Strategy
Everyone's racing to deploy AI agents. Most will waste millions. The question isn't 'how do we use more AI?' - it's 'how do we use AI sustainably?'
-
AI Observability is Expensive Voyeurism
The observability market is selling you dashboards to watch your AI fail in high resolution. What you need is controllability.
-
Your Data Team is Building an AI Graveyard
Every transformation in your data pipeline destroys information AI needs. Traditional data engineering is a lossy compression algorithm.
-
Multi-Agent is This Decade's Microservices Mistake
The multi-agent hype will collapse. We learned this lesson with microservices. Distributed systems are hard.
-
The AI POC Trap: Why Your Demo Worked and Production Won't
Your agentic AI POC impressed leadership. Then you tried to scale it. Here's why demos deceive - and what production actually requires.
-
Foundation Models Are a Commodity. Act Accordingly.
Everyone's agonizing over Claude vs GPT vs Gemini. It doesn't matter. The differentiation is moving up the stack.
-
The LLM Evaluation Maturity Model: Where Does Your Team Actually Stand?
A six-level framework for assessing how your organization evaluates LLM outputs. From 'it looks right' to continuous evaluation pipelines with regression det...
-
The EU AI Act Is Here: What Financial Services Firms Need to Know
A practical guide to EU AI Act compliance for banks, insurers, and investment firms. What's required, what's high-risk, and how to prepare before enforcement...
Newsletter
Occasional, and worth the inbox space.
Notes on agent governance, what the regulations actually say, and what we are building. Roughly monthly. Double opt-in, no tracking, and unsubscribing takes one click and asks you nothing.
RSS works too and needs nothing from you · What happens to your address