Longer piece · 20 minutes
Earned autonomy
Why software went first, what the vendors can't write for you, and how an agent earns the right to act alone.
David Reynolds · September 2026
In plain terms. Software is starting to act in your business's name. The big vendors now sell the controls, but only you can fill them in: who is accountable, what good looks like, and what the software may touch. Let each piece of AI earn more freedom against a written standard, and take it back automatically when its results slip.
A word before starting. An agent, in this piece, is a piece of software that carries out a task on the firm's behalf rather than helping a person do it. A copilot drafts your email. An agent reads the ticket, pulls the file, drafts the response, and sends it, or would if you let it. The question this piece is about is the "if you let it".
Software went first, and it shows where this goes
Software engineering is the useful case because it went first, in public, and it measures itself. Two years ago the tool suggested the next line of code. A year ago it wrote the whole function. This year it takes a ticket from the backlog, reads the codebase, writes the change, runs the tests and submits it for a person to review. In software the debate has moved on from whether agents will do the work to how many, and who checks.
One pattern this year tells you where it goes. The makers of the best-known coding agents have started building a single screen that shows every agent at work, including other companies' agents, and treating their own agent as one plug-in among several. They have decided the lasting position is the place where the work is directed, recorded and approved. I think they're right, and not just for software. Firms will run many agents, the best one will keep changing, and what stays put is the place where the firm decides what agents may do and sees what they did.
Software got there first because checking the work is cheap. Tests pass or they don't, in seconds. Everywhere else, checking is the expensive part. A draft advice note, a customer's contract, a claims decision: somebody qualified has to look. That one difference is why you can't buy the enterprise version of that single screen off the shelf. The screen is easy. What makes it useful is knowing whether the work on it is any good, and in your business only your people know that.
The plumbing is settled, and the accountability isn't
The technical plumbing for agents was standardised this year. Three open standards now do three different jobs, and each has a stable owner: one connects an agent to tools and data, one connects agents to each other, and one connects a person's application to an agent. They're sometimes described as competitors. They're the pipes, and they're largely settled.
Now sit in the first meeting with anyone in your firm who carries risk, and listen to what they ask. What did that task cost? What did the agent actually do? Who is it, and on whose authority is it acting? Who approved it? What is allowed to run here at all? Those five questions come up in the first ten minutes with a finance director, a general counsel, an auditor or a regulator, and none of the three standards answers any of them. The standards move work between agents. They don't say who is responsible for it.
What you can rent, and what only you can make
Most of the AI a business uses this year is rented. The models come from a handful of labs. The copilots sit inside software you already pay for. The platforms that run agents come from the big vendors, and the standards above are free. Your competitor has the same subscriptions, and if they don't yet, they can by Friday.
That is fine. Renting is the right answer for almost all of it. The question is what is left when everyone rents the same things, because whatever is left is what a buyer is paying for when they pay a multiple of your earnings rather than the value of your equipment.
Something changed in 2026 that makes the question urgent. Until recently, if you let software act on the firm's behalf, the controls around it were your problem to build. This year the largest enterprise vendors built the controls into their platforms. An agent can be given an identity, the way an employee has a login. Every call it makes to a model or a tool can be checked against a policy before it goes through. An approval step can be placed in front of anything consequential. Its work can be traced end to end and tested continuously against a standard. Microsoft announced most of this in the first half of the year, and the others are close behind.
Each of those is a mechanism, and each needs something filled in before it does anything. The identity needs the name of the person in your firm who answers for what that agent does. The checkpoint needs a policy, written by you, saying what this agent may and may not touch. The approval step needs a rule for what counts as consequential in your business. The test runner needs tests, and the tests are made of your cases: real files from your own work, with the right answer marked by someone qualified to know it.
So the line between what you rent and what you own has been drawn for you, and it falls in the same place in every firm. On the rented side sit the models, the platforms, the plumbing and the enforcement. On your side sit four things that can only be made from your own operation. (The piece on value adds a fifth, your customers' trust, which you had before AI.)
| Anyone can rent this | Only you can make this |
|---|---|
| Models · copilots · agent platforms · the connection standards · agent identity · policy checks · approval steps · tracing · test runners | Who answers for each agent and what it may touch · your cases with the right answers marked · how much each agent may do without asking, and what it must show to do more · your decisions, recorded alongside your facts |
Each of these is a document, and each can only be written by people who know your business. That is why they are worth something. A competitor can buy your platform tomorrow. They can't buy two years of your professionals marking cases, or your record of who approved what and why. For an owner or a board the practical test is short: take any AI initiative on this year's plan and ask which side of the table the money falls. If all of it is on the left, you are buying capability everyone will have. If some of it lands on the right, you are building something a buyer's diligence team can read and price. The layer you own draws this for a services business, end to end.
One page per agent
The first thing you own is a record, and it is one page. For every agent that exists in the firm, the page says ten things.
What it is and which part of the business it serves. What it may read, call and write, with anything not listed denied. What it may look up and what it must never infer, such as a protected characteristic from an address. How much it is allowed to do on its own, on a four-step scale this piece comes to shortly, held separately for each country or market it operates in. The named person who supervises it, and it is a person, never a team, because a team can't be accountable and everyone in the room knows it. Which model does which step. The set of test cases it is graded against. What it costs per run, how often it raises an exception, and how far its results have drifted from what it scored when it was last approved. Its history: every change of permission, with the date and who decided. And its flags: whether clients are told, where the data stays, how long records are kept.
That is about fifty lines. With it, you can say what is acting in the company's name, on whose authority, how well it works and what it costs, and you can change any of those within a week. Few firms can say that today, because nothing in the plumbing asks for it. I call this the register, because that is what boards already call a document of this kind, and How I run it carries a template for it.
Two details matter. The named supervisor is different from the "sponsor" some platforms now attach to every agent, which is an identity field: it says who owns the account, not who answers for what the agent does. And any change to an agent's permissions, whether a new tool, new data or a wider remit, is a new version that goes back to Observe and earns its way up again. A new model is different: it's re-run against the test set, and the agent keeps its grade if it passes. The model will change more often than the register does.
Line the pages up, agents down the side and markets across the top, with each agent's level in each cell, and you have the one picture a board should see every quarter: what is running, where, and how much it is trusted to do. That is a spreadsheet, and it should look like one.
Your own cases, with the right answers marked
The second thing you own is a test set. For each workflow, in each market, a set of real files from your own work with the right outcome marked by the professionals who do the work: this advice note was correct, this claims decision was wrong and here is why. It's versioned, owned by you, and never handed to the vendors whose models it tests.
Every change to an agent, whether a new model, a new prompt, a new source of data or a new version, is run against the set before the agent is allowed to keep its level, and the score goes on its page. Vendors now ship good tools for running tests like this continuously, and it's fine to use them to run the set. They can't write it.
There are three commercial reasons to fund this, and only one of them is about regulators. It makes changing model vendor nearly free: when something better or cheaper appears you re-run the set and know within a day whether to switch, which is the portability the piece on value said an investor pays for. It turns a claimed productivity gain into a measured one, which is the first step to banking it. And it's the defensible answer to a regulator, an insurer or a buyer who asks how you know the thing works. "The vendor says so" isn't an answer any of them will accept.
None of this is special to professional work. A retailer's pricing agent, an insurer's claims triage and a manufacturer's demand forecast each need the same thing: a set of past cases where someone who knows the work marked what the right call was.
Autonomy is a grade an agent earns, and can lose
Autonomy, here, is a grade. It is held per agent and per market, it is earned against criteria that were published before anyone started measuring, and it is lost automatically. There are four grades, and between them sit the gates, which are the published criteria for moving up and the rule for moving down.
| Grade | What the agent does | How it moves up |
|---|---|---|
| Observe | Works the live queue beside the person; its output is withheld and compared | Agreement with the human decision, over enough volume to mean something |
| Recommend | Drafts; the person decides | Acceptance rate, and an analysis of every case where the person overrode it |
| Prepare | Completes the work; a person approves without material edit | Sustained approval without edits, and a clean run against the test set |
| Act | Acts alone, with a supervisor's signature and a way to reverse it | This is the top. It is held, never granted permanently |
Demotion is the one move in the other direction, and it is automatic. Drift against the test set, a breach of the exception threshold, a move into a new market, or any change to the agent's permissions sends it back to Observe. Publish the demotions as well as the promotions, or nobody will trust the promotions.
The most useful thing in the whole mechanism is the override analysis at the gate out of Recommend. When your people disagree with the agent systematically, the agent is wrong in a specific, fixable way. When they disagree at random, the finding is that your own people disagree with each other, and that is often the more valuable thing to learn about the firm.
Step back from the individual agent and the same idea gives you a picture of the whole firm. Along the bottom, how much your agents are allowed to do without asking. Up the side, how much evidence you hold that they should be. Most firms are in the bottom left, piloting: little autonomy, little evidence, and nothing wrong with that as a place to start. The dangerous firms are in the bottom right, assuming: autonomy has been granted, usually by default when someone switched a feature on, and there is no evidence behind it. The place to be is the top right, earned, and the path there goes up before it goes across. Get the evidence, then grant the autonomy.
For a board, the point is that a grade earned against published criteria is a control an auditor can test and a buyer can check. Autonomy that was assumed is a risk nobody has measured.
Keep the decisions as well as the facts
The fourth thing you own is what the business knows, arranged so an agent can use it. You already have a data estate, so this is a view over the data you already hold, shaped for the first workflow you want an agent to run: which clients exist, what you've done for them, who works with them, what rules apply, which documents matter.
Two pieces of advice, both learned the expensive way. Don't fund an enterprise-wide data model. They usually fail, and you find out after years of spending. Build only what the first workflow needs, which usually turns out to be a handful of kinds of thing, and let each earn its place by being necessary to run that one workflow end to end. The hard part is recognising that the same client in three systems, with three identifiers and two different ideas of who owns it, is one client. Whether you can bring one client together from the systems they're scattered across inside eight weeks is the test of whether any of this is real.
And the design choice that matters most: store decisions as things in their own right. Who signed, on what evidence, with what reasoning. Almost every data estate stores facts and throws the decisions away in email. An agent can't be taught a decision nobody wrote down, and a record of decisions gets more useful the longer you keep it. It's what makes the second agent cheaper to build than the first.
The cheapest model that passes, and how much left the building
Not every step deserves the most expensive model. Steps that carry judgement, where being wrong is costly and volume is low, go to the best model you can rent. High-volume extraction and classification, repetitive and checkable, goes to something cheap or something that runs on your own machines. Anything touching a market with data-residency rules goes to whatever is contractually permitted there, and that constraint beats cost every time.
Choose per step, never per workflow. For the routine steps that are most of the volume, ten times cheaper and two per cent worse is an easy decision, provided your test set exists to prove the two per cent. Choosing per step also means that when a cheaper model passes your tests, you can move to it that week.
The same rule is the privacy policy that can be audited. What a regulator, a client and your general counsel want to know is how much of the client's material left your control to get the answer. So measure it: the volume of client material that left the building, per case. A step that a small local model can pass sends nothing anywhere. That number is more defensible than any policy statement, because it can be checked.
Cost per outcome, reconciled with finance, per workflow, per market, is what turns "AI saves us money" from an opinion into arithmetic. Sometimes the arithmetic says the humans are cheaper, and the right response is to say so and not deploy there.
Start on a Monday
The first three months are mostly publishing rather than building. Month one: a census of every platform and agent already in use, which will be longer than anyone expects, and a baseline agreed with finance, with no platform spend. Month two: the first agent, built to a reusable template and running at Observe from its first day, and the register live. Month three: promotion to Recommend only where the numbers support it, and a first kill list published beside the first results. The detail, for a small team and for a large business, is in How I run it.
Whoever leads it has to own it
Boards are creating a role for this, under one title or another, and the title matters less than what the role owns. It needs two things: ownership of the four documents in this piece (the register, the test set, the gate criteria and the decision record), with the authority to demote an agent the business would rather keep running, and one number to be measured on, the banked column, agreed with the finance director. Documents without the number are governance with nothing to show. The number without the documents is a claim nobody can check. Give the role a seat where investment is decided, and a goal of making itself unnecessary within three years.
What to do with this
If you run the business. Ask for the census. Then ask for one page, for one agent, this month. If nobody can fill in the supervisor line with a name, you have found the first thing to fix.
If you sit on the board. Ask which grade each agent in the business holds and what evidence put it there. If the answer is that nobody has thought of it in those terms, the firm is in the bottom right of the picture and doesn't know it.
If you invest. Ask for the register and the test set in diligence. A firm that has them can test a new model in a day and can prove its AI gains. A firm that has neither is renting its AI estate from someone who can reprice it.
What this rests on
The description of coding agents and the single-screen product is from the public record in 2025 and 2026. The three standards and their owners are public. The Microsoft controls are from Microsoft's own 2026 publications on agent identity, policy enforcement, approval gates and evaluation. The four grades, the gates, the automatic demotion, the override analysis and the register's ten fields are mine; versions of the ladder now appear elsewhere, including AWS's "graduated autonomy" and the Knight Institute's levels of autonomy for agents. What is mine is the specific mechanism above: criteria published before measuring, test sets the firm owns, and grades designed so that a person keeps producing the judgement the agent is measured against. The ninety-day sequence is how I would run it.
The strongest objection
"Staged autonomy slows everything down. By the time an agent has earned Act, the competitor who just switched it on is two years ahead."
I don't accept it. The competitor who switched it on is ahead in the claimed column and nowhere in the banked one, and carrying a risk nobody has measured. The gates don't slow down the work; they slow down the granting of permission, and Observe and Recommend can be running on the first day. What the staged approach costs is a few months per agent per market. What it buys is a control environment an auditor can test and a buyer will pay for. The firm that skipped it will build it later, under a regulator's timetable, at a worse price.
Test it against your own business: Where you are →