An AI playbook · for the person who owns it
How I run it
The pieces say why. This is my playbook for the person who owns AI in a business, whatever the title: chief AI officer, chief digital officer, operating partner or founder. Eleven tools, in the order I use them. Most take an afternoon. Pick your size and the page adjusts.
Under about fifty people. One person owns this, alongside the day job.
What changes at your size
- Who owns it
- One person, alongside the day job, with a sponsor on the leadership team.
- How often you look
- Monthly: the register, the two columns and one scorecard.
- Where the record lives
- A shared spreadsheet and one folder. Live in a week.
- Where you start
- One workflow, thirty test cases, no new spend.
Anything marked with a blue dot changes with your size.
Aim it
1
The four questions
One workflow at a time. An hour with the two or three people who do it.
Point AI at the wrong part of the job and the saving is small. The fourth question is the one people skip, and it's usually the one that changes the answer.
- What is the job to be done? Not the task in the process map: what the person is actually trying to achieve.
- What pain does it remove, for the person doing it and for the customer waiting on it?
- What gain does it create: time, quality, revenue, or something the customer couldn't have before?
- What happens five minutes before it, and five minutes after? Time the whole thing end to end, not just the work.
The reasoning is in Where the time actually goes →
Take it with you: Working files on GitHub
Put in your own numbers for one workflow.
The work is 6% of the time. The prize is almost certainly in the waiting before and after it, not in the work.
Starts with the contract renewal from the piece: four days of work, seventy end to end.
2
Delete, automate, augment
Straight after the four questions, step by step, with the same people in the room.
Automating a step that should have been deleted is the most expensive way to keep it. It now has a vendor and a maintenance cost.
- List every step in the workflow, including the waiting and the rework.
- For each step ask what happens five minutes before and after it, then put it in one of three piles.
- Delete first. Report what you deleted before you report anything you automated.
The reasoning is in If the AI becomes the business →
Take it with you: Working files on GitHub
Delete
Exists because of a system boundary, an old incident or a report nobody reads.
- Re-keying usage figures from one system into another
- A sign-off nobody remembers the reason for
Automate
Routine, checkable and worth doing. Starts at Observe (tool 6).
- Gathering usage before anyone asks
- Flagging the clauses this customer always changes
Augment
Carries judgement, or is how people learn (tool 8). Software drafts, a person decides.
- Pricing the renewal
- Agreeing the final terms
3
Stages and gates
Before anything is built. One page per initiative, and the owner decides at each gate.
Agreeing first what would make you stop is what makes stopping possible. AI makes every stage faster. It doesn't remove a gate, and it doesn't mean jumping straight to scale.
- Write down, before the work starts, the result that would make you stop.
- Take a baseline before you build, so there's something to measure against.
- At every gate, decide in the room: carry on or stop, on the measured evidence.
- Publish what you stopped beside what you scaled. It's what makes the wins believable.
The reasoning is in Where the time actually goes →
Take it with you: Working files on GitHub
Written down before the work starts: what would make us stop
- 1
Ideation
Shape it until someone would pay for it, or the people who'd use it say they would.
Go or stop - 2
Design
Test it with those people, on their own work, before building anything properly.
Go or stop - 3
Execution
Build it, and measure it against the baseline taken first.
Go or stop - 4
Acceleration
Make it faster and cheaper without losing the quality you measured.
Go or stop - 5
Scaling
Take it to the next team or market, configured rather than rebuilt.
Prove it
4
The two columns
Kept by whoever holds the budget. Looked at monthly.
Time saved isn't money until someone banks it. A buyer's diligence team discounts the first column heavily and pays for the second.
- Claimed: every saving a tool is said to produce, with where the estimate came from.
- Banked: only what passed one of two tests in the same quarter. The freed time went to named work that earns, or the cost left the accounts.
- A line moves across when someone can point to it in the accounts. Never on a forecast.
- The ratio between the columns is the finding. Report both together.
The reasoning is in If the AI becomes the business →
Take it with you: Two columns template (CSV) · Working files on GitHub
| Claimed | Banked | |
|---|---|---|
| Drafting tool, client team | 4 hrs a person a week, 20 people | Nothing yet. No named work has taken the hours |
| Invoice matching automated | 1.5 people's worth of effort | 1 role not refilled in Q3 |
| First review, agent drafting | 30% faster per file | 30% more files per reviewer at the same fixed fee, showing in margin since Q3 |
| Meeting summaries | 2 hrs a person a week | Nothing, and nothing planned. Leave it claimed |
Most businesses only keep the left-hand column. The first time they fill in the right-hand one, it's usually much smaller.
5
The test set
One per workflow. Start with thirty cases.
The real answer to “how do you know it works?”. It's also what lets you test a new model in a day rather than a quarter.
- Ask the people who do the work best to pick real cases from the last two years: clear passes, clear fails, and the awkward ones in between.
- Mark the right outcome, and for the fails, why. Strip anything that mustn't leave the business.
- Version it. Never share it with the vendors whose models it tests. Their tools may run it; they never write it.
- Run it on every change to an agent (model, prompt, data, version) before it keeps its grade.
- Refresh a share every quarter. A set nobody has touched in a year is measuring last year's business.
The reasoning is in Earned autonomy →
Take it with you: Working files on GitHub
- Clear passes · 13
- Clear fails, with the reason · 10
- The awkward ones in between · 7
6
The grades and the gates
Set the criteria in writing before anything is measured against them. One set for the whole company.
Autonomy is a grade an agent earns against criteria published in advance, and loses automatically. A grade like that is a control an auditor can test. Autonomy that was assumed is a liability nobody has priced.
- Every agent starts at Observe, whatever the vendor says it can do.
- Write the gate for each grade, with your own thresholds, and publish it first.
- Demotion is automatic: drift against the test set, too many exceptions, a new market, or any change to what it may touch.
- Report demotions beside promotions, so the promotions can be trusted.
The reasoning is in Earned autonomy →
Take it with you: Working files on GitHub
Observe
What it does. Works the live queue beside a person. Its output is held back and compared.
To reach Recommend. Agreement with the person's decision above your threshold, over enough cases to mean something. Set the number of cases too.
Back to Observe, automatically: drift against the test set · too many exceptions · a new market · any change to what it may touch or look up. A new model is re-tested before the agent keeps its grade.
7
The register
A shared spreadsheet, one row per agent. Live in a week.
With it you can say what is acting in your name, on whose authority, how well it works and what it costs. If it isn't on the register, it doesn't run.
- Every agent declares ten things. Anything not declared is denied.
- The supervisor is a named person. Never a team, never a role title.
- Any change to what it may touch or look up is a new version. It goes back to Observe. A new model is re-tested before it keeps its grade.
- Once a month, look at the whole register together.
The reasoning is in Earned autonomy →
Take it with you: Register template (CSV) · Register template (Markdown) · Working files on GitHub
An example page
Renewal prep agent · v3
- 1Identity and purpose
- Gathers usage and flags likely changes before each renewal. Commercial team.
- 2May read, call, write
- Reads usage and contract history. Writes a draft pack. Nothing else.
- 3Must never
- Contact a customer. Infer anything about a named person.
- 4Grade
- Recommend
- 5Supervisor
- A named person on the commercial team
- 6Model per step
- Small model for extraction · best available for the clause flags
- 7Test set
- Renewals v2 · 41 cases · last score 94%
- 8Cost and health
- Cost per renewal · exception rate · drift since last promotion
- 9History
- v2 demoted after a pricing change · v3 promoted to Recommend after 6 weeks
- 10Flags
- Customers told · data stays in region · kept for 2 years
| Agent | Grade |
|---|---|
| Renewal prep | R |
| Support triage | P |
| Invoice matching | A |
| Contract drafting | O |
The whole register at a glance, looked at monthly. O Observe · R Recommend · P Prepare · A Act
Keep it
8
The curriculum exercise
Before your next three hires. An hour.
Most automation decisions sort tasks by what the output is worth. The second question is what doing the task taught. Miss it and the saving arrives this year while the cost arrives in five.
- List what a newer person does in a normal month. Ask them, and ask whoever reviews their work.
- Mark each task on two things: what the output is worth, and how much doing it taught.
- Place each task in the grid, and act on the box it lands in.
- Write down which work you're keeping for people because it teaches, even where an agent could do it.
The reasoning is in The bottom was the training programme →
Take it with you: Working files on GitHub
Automate
High value, taught little. Automate it and bank the gain the same quarter.
- Matching invoices to orders
- Chasing missing receipts
Draft and decide
High value, and it was how people learned. Software drafts, a person decides.
- Month-end commentary
- Spotting an odd supplier
Delete
Low value, taught nothing. Delete it. No software needed.
- Re-formatting reports for three audiences
Keep a share
Low value, but it taught. Automate the volume, keep a set share for people, and call it a training budget.
- Checking the routine accruals
9
The model, step by step
Per step, never per workflow. Re-decide twice a year, or when prices move.
Not every step deserves the most expensive model. Ten times cheaper and two per cent worse is an easy decision, provided your test set exists to prove the two per cent.
- Hard, rare, costly-if-wrong steps go to the best model you can rent.
- High-volume, checkable steps go to the cheapest model that passes your test set, including ones that run on your own machines.
- Where data rules apply, they come first, whatever the cost.
- Measure how much customer material left your control, per case. That number can be checked; a policy can't.
- Reconcile cost per outcome with finance. If people are cheaper for a step, say so and don't deploy there.
The reasoning is in Earned autonomy →
Take it with you: Working files on GitHub
Best model you can rent
Rare, hard, costly if wrong
Best model, plus a person
Frequent and costly: this is where the test set earns its keep
Whatever passes
Rare and cheap to get wrong: don't over-think it
Cheapest that passes
High volume, checkable: small or local models
Data rules beat cost, every time.
10
The first ninety days
Almost all of it is publishing rather than building.
What makes the first three months credible is that other people can read what you did. A kill list beside the wins works because other people can check it.
The reasoning is in Earned autonomy →
Take it with you: Working files on GitHub
Month one
- List every AI tool and agent already in use
- Agree the baseline for one workflow
- No new spend
Month two
- First agent, at Observe from day one
- Register live, even with one row
- Gate criteria written down
Month three
- Promote to Recommend if the numbers say so
- First scorecard and first kill list
- Decide: carry on or stop
Published along the way
- The agent standard
- The register
- The gate criteria
- The baseline
- The first scorecard
- The first kill list
One person owns it, alongside their day job, with a named sponsor on the leadership team.
11
Who owns it
Part of someone's job.
The title matters less than what the role owns: four documents and one number. Without the number it governs but delivers little; without the documents it delivers numbers nobody can check.
- It owns four documents: the register, the test set, the gate criteria and the record of decisions.
- It can demote an agent the business would rather keep running.
- It's measured on one number: the banked column, agreed with whoever runs the money.
- In a small company this is usually the founder or the CTO, for a few hours a week. That's enough, as long as it's written down whose it is.
The reasoning is in Earned autonomy →
Take it with you: Working files on GitHub
- The register
- The test set
- The gate criteria
- The record of decisions
Measured on one number: the banked column
Every tool here links to the piece that explains why it works. The whole set is also on GitHub: every tool as markdown files you can use with your own team or your own AI tools, a thirty-point self-check for practitioners, and a Claude plugin that runs five of the exercises with you. github.com/reynolds-uk/ai-playbook →