rotascale

Blog

Nobody Is Making You Govern Your Agents. Do It Anyway.

TL;DR: This category markets almost exclusively to firms with a supervisor, which means it is aimed away from the teams with the largest agent estates and the fastest release cycles. If you are a platform or an AI-native product company, no statute is coming for you soon. Your enterprise buyer’s security questionnaire already did, and there are two other reasons that have nothing to do with anybody’s rules.

The assumption is backwards

The pitch goes: regulated industries must govern their AI, so governance is a compliance purchase, so the buyer is a bank.

Look at where the agents actually are. The largest fleets, the most autonomy, the shortest paths from idea to production, and the widest access to other people’s data are not in banks. They are in software companies, shipping agent features into customer accounts weekly, with the agent holding a credential that reaches across every tenant.

The bank is being careful and slow. You are being fast, which is correct for your business, and it means you accumulated the exposure first.

You are upstream, not exempt

Here is the part that decides this.

When a hospital, an insurer or a wholesale bank puts an agent into production, a meaningful part of that agent is somebody’s platform. The retrieval layer, the orchestration framework, the agent runtime, the vertical SaaS it acts inside. Yours, quite possibly.

Every question their supervisor asks them, they will ask you. Not through a regulator. Through a contract, with a deadline you did not negotiate, in a security review that gates a renewal. That transmission is already happening and it does not require any new law.

So the practical question is not “am I regulated”. It is “can I answer the questions my customers are being asked”, and the companies that can will take those accounts from the ones that cannot.

Three reasons before you get to any of that

Blast radius. Your agent acts inside customer accounts. One credential, many tenants, and the boundary between them is your own code being careful in every path. That is a strong assumption to hold across a codebase several teams ship to weekly. One prompt-injected document in tenant A that reaches tenant B is not a bug report, it is a disclosure obligation and a set of phone calls.

Spend. An autonomous agent with an API key and a loop is an unbounded cost centre until something bounds it. Not just inference. Refunds, credits, third-party calls, compute. The characteristic failure is not one expensive mistake, it is thousands of individually reasonable actions with no ceiling on the total, discovered at month end.

Trust, which is the one that makes money. “What did your agent do in my account last quarter?” is a question you will be asked, and the answer decides renewals above a certain contract size. Right now most companies answer it with a screenshot of an internal dashboard and a promise.

The thing you can do that a bank cannot

Everyone else in this category produces evidence for a filing cabinet. Packs for an examiner, read once, archived.

You can put it in the product.

A per-tenant, sealed, independently verifiable record of what your agent did inside a customer’s account, rendered in your own UI, exported by the customer, checked by their security team without touching your infrastructure. That is a feature you ship, on the same afternoon it stops being a risk you carry.

It is a genuinely unusual position. The compliance artefact and the differentiating feature are the same object. Nobody selling to banks gets to say that.

Why it does not get built

Three objections, and they are all reasonable.

It slows us down. It would, if adopting it meant a runtime or a proxy. It should be a call in your own code before a consequential action, on the stack you already chose, with the check ordered so cheap structural failures are the cheap ones.

We do not know what the limits should be. Correct, and that is the actual work. It is also why the first mode should record rather than refuse: run it against production for a fortnight, find out what your agents are really doing, and set limits from data instead of from a meeting. The first week reliably contains at least one behaviour nobody knew about.

We will do it when a customer asks. By then it is a blocker on a deal with a date on it, and you are building under the worst possible conditions. The version built calmly is better and takes less time.

What to do this quarter

List the agents and name an owner for each. Not the team. A person. The list will be longer than you think and some of it will have no owner at all. That finding is worth the afternoon.

For each consequential action, write the sentence. “This agent may do X, up to Y, on behalf of Z, until W.” If nobody can write it, that is the gap, and no software fixes it.

Turn on recording, refuse nothing. Zero operational risk, immediate visibility, and it produces the evidence for the argument you will need to make internally.

Then move one agent to enforcing. One. The one where a mistake is most expensive. Watch what it refuses for a fortnight before touching anything else.

The bottom line

Governance is being sold as a compliance product to people who are compelled. That framing will keep it out of the hands of the teams that need it most, for another year or two, until a large customer or a bad afternoon forces it.

You are not exempt because nobody is making you. You are early because nobody is making you, and being early is the only time this is cheap.

Newsletter

If this was useful, the next one is too.

Notes on agent governance, what the regulations actually say, and what we are building. Roughly monthly. Double opt-in, no tracking, and unsubscribing takes one click and asks you nothing.

RSS works too and needs nothing from you · What happens to your address