An AI playbook · for the person who owns it

How I run it

The pieces say why. This is my playbook for the person who owns AI in a business, whatever the title: chief AI officer, chief digital officer, operating partner or founder. Eleven tools, in the order I use them. Most take an afternoon. Pick your size and the page adjusts.

I'm working in

Under about fifty people. One person owns this, alongside the day job.

What changes at your size

Who owns it
One person, alongside the day job, with a sponsor on the leadership team.
How often you look
Monthly: the register, the two columns and one scorecard.
Where the record lives
A shared spreadsheet and one folder. Live in a week.
Where you start
One workflow, thirty test cases, no new spend.

Anything marked with a blue dot changes with your size.

Also on GitHub, as markdown files and a Claude plugin →

Aim it

1

The four questions

One workflow at a time. An hour with the two or three people who do it.

Point AI at the wrong part of the job and the saving is small. The fourth question is the one people skip, and it's usually the one that changes the answer.

  1. What is the job to be done? Not the task in the process map: what the person is actually trying to achieve.
  2. What pain does it remove, for the person doing it and for the customer waiting on it?
  3. What gain does it create: time, quality, revenue, or something the customer couldn't have before?
  4. What happens five minutes before it, and five minutes after? Time the whole thing end to end, not just the work.

The reasoning is in Where the time actually goes →

Take it with you: Working files on GitHub

Put in your own numbers for one workflow.

Waiting beforeThe workWaiting and rework after

The work is 6% of the time. The prize is almost certainly in the waiting before and after it, not in the work.

Starts with the contract renewal from the piece: four days of work, seventy end to end.

2

Delete, automate, augment

Straight after the four questions, step by step, with the same people in the room.

Automating a step that should have been deleted is the most expensive way to keep it. It now has a vendor and a maintenance cost.

  1. List every step in the workflow, including the waiting and the rework.
  2. For each step ask what happens five minutes before and after it, then put it in one of three piles.
  3. Delete first. Report what you deleted before you report anything you automated.

The reasoning is in If the AI becomes the business →

Take it with you: Working files on GitHub

Delete

Exists because of a system boundary, an old incident or a report nobody reads.

  • Re-keying usage figures from one system into another
  • A sign-off nobody remembers the reason for

Automate

Routine, checkable and worth doing. Starts at Observe (tool 6).

  • Gathering usage before anyone asks
  • Flagging the clauses this customer always changes

Augment

Carries judgement, or is how people learn (tool 8). Software drafts, a person decides.

  • Pricing the renewal
  • Agreeing the final terms

3

Stages and gates

Before anything is built. One page per initiative, and the owner decides at each gate.

Agreeing first what would make you stop is what makes stopping possible. AI makes every stage faster. It doesn't remove a gate, and it doesn't mean jumping straight to scale.

  1. Write down, before the work starts, the result that would make you stop.
  2. Take a baseline before you build, so there's something to measure against.
  3. At every gate, decide in the room: carry on or stop, on the measured evidence.
  4. Publish what you stopped beside what you scaled. It's what makes the wins believable.

The reasoning is in Where the time actually goes →

Take it with you: Working files on GitHub

Written down before the work starts: what would make us stop

  1. 1

    Ideation

    Shape it until someone would pay for it, or the people who'd use it say they would.

    Go or stop
  2. 2

    Design

    Test it with those people, on their own work, before building anything properly.

    Go or stop
  3. 3

    Execution

    Build it, and measure it against the baseline taken first.

    Go or stop
  4. 4

    Acceleration

    Make it faster and cheaper without losing the quality you measured.

    Go or stop
  5. 5

    Scaling

    Take it to the next team or market, configured rather than rebuilt.

Prove it

4

The two columns

Kept by whoever holds the budget. Looked at monthly.

Time saved isn't money until someone banks it. A buyer's diligence team discounts the first column heavily and pays for the second.

  1. Claimed: every saving a tool is said to produce, with where the estimate came from.
  2. Banked: only what passed one of two tests in the same quarter. The freed time went to named work that earns, or the cost left the accounts.
  3. A line moves across when someone can point to it in the accounts. Never on a forecast.
  4. The ratio between the columns is the finding. Report both together.

The reasoning is in If the AI becomes the business →

Take it with you: Two columns template (CSV) · Working files on GitHub

ClaimedBanked
Drafting tool, client team4 hrs a person a week, 20 peopleNothing yet. No named work has taken the hours
Invoice matching automated1.5 people's worth of effort1 role not refilled in Q3
First review, agent drafting30% faster per file30% more files per reviewer at the same fixed fee, showing in margin since Q3
Meeting summaries2 hrs a person a weekNothing, and nothing planned. Leave it claimed

Most businesses only keep the left-hand column. The first time they fill in the right-hand one, it's usually much smaller.

5

The test set

One per workflow. Start with thirty cases.

The real answer to “how do you know it works?”. It's also what lets you test a new model in a day rather than a quarter.

  1. Ask the people who do the work best to pick real cases from the last two years: clear passes, clear fails, and the awkward ones in between.
  2. Mark the right outcome, and for the fails, why. Strip anything that mustn't leave the business.
  3. Version it. Never share it with the vendors whose models it tests. Their tools may run it; they never write it.
  4. Run it on every change to an agent (model, prompt, data, version) before it keeps its grade.
  5. Refresh a share every quarter. A set nobody has touched in a year is measuring last year's business.

The reasoning is in Earned autonomy →

Take it with you: Working files on GitHub

  • Clear passes · 13
  • Clear fails, with the reason · 10
  • The awkward ones in between · 7

6

The grades and the gates

Set the criteria in writing before anything is measured against them. One set for the whole company.

Autonomy is a grade an agent earns against criteria published in advance, and loses automatically. A grade like that is a control an auditor can test. Autonomy that was assumed is a liability nobody has priced.

  1. Every agent starts at Observe, whatever the vendor says it can do.
  2. Write the gate for each grade, with your own thresholds, and publish it first.
  3. Demotion is automatic: drift against the test set, too many exceptions, a new market, or any change to what it may touch.
  4. Report demotions beside promotions, so the promotions can be trusted.

The reasoning is in Earned autonomy →

Take it with you: Working files on GitHub

Observe

What it does. Works the live queue beside a person. Its output is held back and compared.

To reach Recommend. Agreement with the person's decision above your threshold, over enough cases to mean something. Set the number of cases too.

Back to Observe, automatically: drift against the test set · too many exceptions · a new market · any change to what it may touch or look up. A new model is re-tested before the agent keeps its grade.

7

The register

A shared spreadsheet, one row per agent. Live in a week.

With it you can say what is acting in your name, on whose authority, how well it works and what it costs. If it isn't on the register, it doesn't run.

  1. Every agent declares ten things. Anything not declared is denied.
  2. The supervisor is a named person. Never a team, never a role title.
  3. Any change to what it may touch or look up is a new version. It goes back to Observe. A new model is re-tested before it keeps its grade.
  4. Once a month, look at the whole register together.

The reasoning is in Earned autonomy →

Take it with you: Register template (CSV) · Register template (Markdown) · Working files on GitHub

An example page

Renewal prep agent · v3

1Identity and purpose
Gathers usage and flags likely changes before each renewal. Commercial team.
2May read, call, write
Reads usage and contract history. Writes a draft pack. Nothing else.
3Must never
Contact a customer. Infer anything about a named person.
4Grade
Recommend
5Supervisor
A named person on the commercial team
6Model per step
Small model for extraction · best available for the clause flags
7Test set
Renewals v2 · 41 cases · last score 94%
8Cost and health
Cost per renewal · exception rate · drift since last promotion
9History
v2 demoted after a pricing change · v3 promoted to Recommend after 6 weeks
10Flags
Customers told · data stays in region · kept for 2 years
AgentGrade
Renewal prepR
Support triageP
Invoice matchingA
Contract draftingO

The whole register at a glance, looked at monthly. O Observe · R Recommend · P Prepare · A Act

Keep it

8

The curriculum exercise

Before your next three hires. An hour.

Most automation decisions sort tasks by what the output is worth. The second question is what doing the task taught. Miss it and the saving arrives this year while the cost arrives in five.

  1. List what a newer person does in a normal month. Ask them, and ask whoever reviews their work.
  2. Mark each task on two things: what the output is worth, and how much doing it taught.
  3. Place each task in the grid, and act on the box it lands in.
  4. Write down which work you're keeping for people because it teaches, even where an agent could do it.

The reasoning is in The bottom was the training programme →

Take it with you: Working files on GitHub

What the output is worth →

Automate

High value, taught little. Automate it and bank the gain the same quarter.

  • Matching invoices to orders
  • Chasing missing receipts

Draft and decide

High value, and it was how people learned. Software drafts, a person decides.

  • Month-end commentary
  • Spotting an odd supplier

Delete

Low value, taught nothing. Delete it. No software needed.

  • Re-formatting reports for three audiences

Keep a share

Low value, but it taught. Automate the volume, keep a set share for people, and call it a training budget.

  • Checking the routine accruals
What doing it taught →

9

The model, step by step

Per step, never per workflow. Re-decide twice a year, or when prices move.

Not every step deserves the most expensive model. Ten times cheaper and two per cent worse is an easy decision, provided your test set exists to prove the two per cent.

  1. Hard, rare, costly-if-wrong steps go to the best model you can rent.
  2. High-volume, checkable steps go to the cheapest model that passes your test set, including ones that run on your own machines.
  3. Where data rules apply, they come first, whatever the cost.
  4. Measure how much customer material left your control, per case. That number can be checked; a policy can't.
  5. Reconcile cost per outcome with finance. If people are cheaper for a step, say so and don't deploy there.

The reasoning is in Earned autonomy →

Take it with you: Working files on GitHub

Cost of being wrong →

Best model you can rent

Rare, hard, costly if wrong

Best model, plus a person

Frequent and costly: this is where the test set earns its keep

Whatever passes

Rare and cheap to get wrong: don't over-think it

Cheapest that passes

High volume, checkable: small or local models

Volume →

Data rules beat cost, every time.

10

The first ninety days

Almost all of it is publishing rather than building.

What makes the first three months credible is that other people can read what you did. A kill list beside the wins works because other people can check it.

The reasoning is in Earned autonomy →

Take it with you: Working files on GitHub

Month one

  • List every AI tool and agent already in use
  • Agree the baseline for one workflow
  • No new spend

Month two

  • First agent, at Observe from day one
  • Register live, even with one row
  • Gate criteria written down

Month three

  • Promote to Recommend if the numbers say so
  • First scorecard and first kill list
  • Decide: carry on or stop

Published along the way

  • The agent standard
  • The register
  • The gate criteria
  • The baseline
  • The first scorecard
  • The first kill list

One person owns it, alongside their day job, with a named sponsor on the leadership team.

11

Who owns it

Part of someone's job.

The title matters less than what the role owns: four documents and one number. Without the number it governs but delivers little; without the documents it delivers numbers nobody can check.

  1. It owns four documents: the register, the test set, the gate criteria and the record of decisions.
  2. It can demote an agent the business would rather keep running.
  3. It's measured on one number: the banked column, agreed with whoever runs the money.
  4. In a small company this is usually the founder or the CTO, for a few hours a week. That's enough, as long as it's written down whose it is.

The reasoning is in Earned autonomy →

Take it with you: Working files on GitHub

  • The register
  • The test set
  • The gate criteria
  • The record of decisions

Measured on one number: the banked column

Every tool here links to the piece that explains why it works. The whole set is also on GitHub: every tool as markdown files you can use with your own team or your own AI tools, a thirty-point self-check for practitioners, and a Claude plugin that runs five of the exercises with you. github.com/reynolds-uk/ai-playbook →

How I run it: an AI playbook · AI, and what a business is worth