How to Turn an SOP Into an AI Agent

Somewhere in your Google Drive there’s a document called something like “How we handle proposals”. Someone wrote it so a new hire could do the job without asking. It lists the steps, says what to check, and says when to ask the owner.

That document is most of the spec for an AI agent. Deciding exactly what it should do, and when it should stop, is already on the page.

The part that matters most usually isn’t. One of my clients, a three-person design studio, had every proposal waiting on the founder, with a median delay of four days. The founder was the bottleneck because he decided which jobs were standard and which weren’t, using rules he’d never written down. Writing them down took most of the three weeks of part-time work the build needed. Now proposals under $5,000 go out in under three hours.

Here’s how to go from a written SOP to something that runs: eight steps, from choosing the right SOP to widening what it’s allowed to do on its own. The studio is the example throughout.

1. Why this works now

An SOP is a list of steps with a few decisions in between. Software could always follow the steps. It couldn’t make the decisions, because SOP decisions are written for people: “check the scope is standard”, “use judgement on rush jobs”.

Language models can now read a step like that and apply it consistently, provided the rule is written clearly and they have what a person would have. They can also return the decision in a fixed format, so an automation can act on it without anyone reading it first.

So turning “a document a person follows” into “a process that runs” is now mostly a rewriting job.

Here’s where I’d push back on the usual advice. “Turn your SOP into an agent” makes people picture one AI that reads the document and works out the whole job. That’s the wrong build for most SOPs. Most of their steps are fixed: look up the price, check the total, send the email. Those should be plain automation. The model belongs at the few steps that need judgement.

I keep a library of 50 client builds, wins and failures. In 24 of the 26 wins, what produced the result was a check, a permission, a written file, or a clear decision about who decides. The quality of the Skill itself was the main driver in two.

Anthropic’s engineering team draws a similar line in Building effective agents. They call systems that follow “predefined code paths” workflows, and systems where the model directs its own process agents, and they recommend “finding the simplest solution possible, and only increasing complexity when needed.” An SOP is a predefined path. Start there.

2. The workflow, step by step

Step 1: Pick the right SOP, and the right kind of build

Your first SOP should pass four tests.

It runs often. Weekly at least. A process that runs twice a year isn’t worth automating, and it won’t give you enough cases to test against.

The rules are mostly written down, or could be. If the honest answer to “how do you decide?” is “I just know”, plan for that to be the main work, as it was for the studio.

A mistake is cheap and easy to undo. A wrong tag on a ticket is fine. A wrong $8,000 quote is a bad first project.

The inputs are already text. If the process starts with a phone call or a paper form, getting that into text is a separate project.

Then decide what it should become. My rule of thumb: if the work turns up now and then, a saved prompt is enough. If it happens the same way every time and needs your standard, make it a Skill you run yourself. It only needs to be an agent if it has to start without you. And chain several of those into a workflow last, once each piece works on its own. I explained how I decide between the four in a Bionic Business issue, When To Use A Prompt, Skill, An Agent, Or Yourself.

Plenty of SOPs never need an agent. A solo developer I work with turned his pre-handoff checklist into a saved 22-point prompt that he runs by hand on every build. Across 31 handoffs, it caught 74 issues, and bug reports after handoff fell from 4.1 per project to 0.8. A virtual assistant working alone built three saved prompts, one of which turns each new client’s intake call into a set of SOPs. Her onboarding time per client fell from about 9 hours to 3.5. Neither of them has an agent.

Step 2: Rewrite it in five columns

Most SOPs are prose or loose bullets. Rewrite every step into the same five columns: what comes in, what happens, what comes out, the rule for any decision, and who or what decides.

Here’s the studio’s proposal process in those columns, reconstructed from its build record:

#InWhat happensOutRuleWho decides
1A qualified enquiry and the call notesTurn the notes into a list of deliverables, in the studio’s own vocabularyDeliverables listNoneModel
2Deliverables listCheck for anything that makes it customStandard or customRush timelines and multi-brand work are custom, per the custom-scope fileModel, reading the file
3Deliverables listPrice each item from the published tiersPriced listEvery number must come from the pricing fileModel, then a rule that checks every figure
4Priced listPut it into one of six approved proposal structuresDraft proposalApproved structures onlyModel
5Draft proposalCheck the totalSend or hold$5,000 or more cannot sendRule
6Custom or held draftsReview, edit, sendSent proposalAll custom scope and anything over $5,000Founder
7Sent proposalsPull a weekly sample of fiveReviewed sampleNoneStudio manager

Doing this rewrite shows you two things straight away. Three of the seven rows need no model at all, and the model rows lean on files and fixed checks. And every vague phrase in the original (“check it’s standard”) has to become a rule someone could check.

If the SOP is long, a model can do the first pass. Paste it in with something like this:

Here is a process I follow. Rewrite it as a table with these columns: step number, what comes in, what happens, what comes out, the rule for any decision, and who decides (a person, a fixed rule, or a judgement call).

Mark every point where I decide based on experience or gut feel rather than a written rule. For each one, ask me the question you'd need answered to write the rule down. Don't guess the answer.

[paste the SOP]

The questions it asks are the real output. Each one is a rule that lives in your head.

Step 3: Sort every row into judgement, fact or limit

My case library is organised around this rule, and it explains most of the failures in it.

A judgement goes to the model: a prompt, a Skill or an agent. Translating call notes into deliverables is a judgement.

A fact goes in a file the model reads every time it runs. Prices, criteria, terms, exclusions. The studio’s pricing file is the only place any number can come from.

A limit goes in a workflow step that passes or fails with no opinion. “$5,000 or more can’t send” is a limit. So is “a figure not in the pricing file fails the build”.

Most of the failures I’ve documented are a fact or a limit that somebody wrote into the agent’s instructions and hoped it would follow. A model can ignore an instruction. A workflow step can’t.

Step 4: Write down what isn’t written down

Every row marked “gut feel” in step 2 needs a written rule before you build.

This is where the time goes. For the studio, it was a file listing what sends a proposal to the founder: rush timelines, work across several brands, anything above $5,000. The founder had been applying those rules by instinct. Writing them down took most of the three weeks of part-time work the build needed.

If you can’t write the rule, keep that step with a person. That’s a fine outcome. You’ve still automated everything around it.

Step 5: Map each row to a tool and build version one

Rule rows go to whatever automation you already use: Make, Zapier, n8n, or a short script. Model rows become one prompt each, doing one job, with the relevant facts files attached. Person rows become a message to wherever you’ll see it, with the draft and a way to approve.

Keep the facts in files the build reads at run time, never pasted into the prompt. When the price list changes, you update one file and every proposal after that uses it.

For writing the prompts themselves, the structure in my guide to Claude system prompts works with any model. If the SOP depends on a lot of background knowledge, how to train an AI agent on company data covers how to give it that context.

Step 6: Test it on past cases

Before it sees a live case, run it on 20 to 40 past ones where you know what the right outcome was. Pull them from your sent folder, your CRM or your log.

Compare its output with what you actually did, row by row. Count the matches. Then read every mismatch. Some will be the model’s mistake, some will be yours, and some will show you a rule that was never in the SOP at all.

Step 7: Run it in shadow mode

Next, run it on live work alongside the person who normally does it, with sending switched off. The person works as usual. The build does the same case in the background and logs what it would have done.

The studio ran 34 proposals this way over about seven weeks. The founder’s version and the agent’s matched on 31. All three disagreements were projects above $5,000 with custom scope, the class the founder had already decided to keep for himself. Nobody had to argue about whether the agent was ready, because there was a number.

Shadow mode finds what you didn’t know to test. A four-person IT consultancy I work with ran its enquiry routing in shadow for about three and a half weeks. The agent matched the owner on 36 of 41 enquiries. Four of the five disagreements turned out to be current clients of the firm’s biggest referral partner. The owner had been sending those to himself for three years, from memory, and the rule existed nowhere. Adding it took twenty minutes once the disagreements made it visible. The next 12 enquiries matched 12 for 12. That’s a small test set, and it checks routing only, but it moved about five hours a week off the owner at a point when he was the limit on sales.

The disagreements are what you’re testing for. Don’t tune them away. Read each one and ask what rule it points to.

One detail from that build worth copying: shadow mode was a setting in the workflow, not an instruction in the prompt. Nothing could leak out during testing, because the workflow held it in shadow, whatever the prompt said.

Step 8: Go live on one slice, then widen it on evidence

Give the build authority over the safest slice first, in writing, and keep everything else with a person.

The studio’s agent was allowed to send proposals under $5,000 without review, and nothing else. The next step is written down too: the limit rises to $8,000 after 40 proposals in a row with no scope dispute and no override from the founder. The first dispute sends it back to review.

Expect to review every output for the first couple of weeks. Once the output has been clean for six to eight weeks, move to a weekly sample. The studio’s manager reviews five sent proposals a week.

3. What it replaces

For the studio: proposals under $5,000 went from a four-day wait to under three hours, and they make up 71% of the volume. The close rate on that tier rose from 41% to 56% in the first four weeks, which meant nine extra projects at an average of $3,100. Four weeks is a short window, and September has brought a seasonal lift in this studio before, so that number has to hold for longer before it means much.

For the IT consultancy: about five hours a week of the owner’s time.

For the solo examples: about five and a half hours of unpaid onboarding per new client for the virtual assistant, and fewer post-handoff bug reports for the developer.

4. What it needs to work

The SOP itself, up to date. If people follow a newer version in their heads, find that out first, because the build will follow the paper.

The facts as files: prices, criteria, exclusions, terms, each with a date and someone who maintains it.

20 to 40 past cases with known outcomes, for step 6.

A named owner. This sounds like admin. Across my case library, 14 of the 18 builds that cost money had no recorded owner when they failed, and 25 of the 26 wins had one before launch.

5. The number to watch

Before launch, the match rate: how often the build made the same call as the person on the same cases. The studio’s was 31 of 34. The IT consultancy’s was 36 of 41, then 12 of 12 after the fix. Don’t go live until you’ve read every mismatch and either fixed the cause or decided it’s acceptable.

After launch, the override rate: how often a person changes what the build did. It should fall as you fix rules, then hold steady. If it starts rising again, something changed: the inputs, the model or the business.

6. Where it breaks

Knowledge that lives in someone’s head. The referral-partner rule above sat in one person’s memory for three years. Shadow mode is where these surface, which is why you don’t skip it.

Old examples carry old mistakes. Another client, a $2.4M productized SEO service, built a quoting agent that used twelve past proposals as examples of good style. Two of them included a 15% first-month discount that a former salesperson had given without permission. The agent learned it. Over ten weeks, 19 of 63 quotes carried the discount, 11 of those clients signed, and the business gave away $6,845. Nobody flagged it, because the discount was in the business’s own voice and the deals looked good. Compare the studio, whose prices can only come from the pricing file and where any other figure fails the build. If your SOP’s past outputs will be used as examples, clean them first, and keep anything with money in it in a file the build reads.

AI-written SOPs need a sign-off. The virtual assistant’s clients approve their SOP set in writing before anything runs on it. That step caught three wrong assumptions in her first two onboardings.

SOPs that drift. The business changes a policy in a meeting and nobody updates the document. The build keeps applying the old policy, perfectly. Whoever changes a policy changes the file.

Model updates. The same prompt can behave differently after the provider updates the model. Keep your past cases and run them again after any change.

Pick one SOP this week, rewrite it in the five columns, and count the rows that need a model. If it’s one or two, you can build it with the tools you already pay for.

You’re subscribed. Read this week’s issue →

Get the workflows as I build them

Every week I share the agents and systems for growing your online business: the full build, the data they need, and exactly how you can run the same. Thousands of operators read Bionic Business. You should, too. No hype, no fluff. Only what's working now.

Free. One email a week. Unsubscribe whenever you like. Read this week’s issue first

Sam Woods

Written by

Sam Woods

Fractional Chief AI Officer · Founder, Stimulead and Daring Robot

Sam started with machine learning in 2016 and generative AI in 2019, writing production prompts before the practice had a name. He has advised and trained Fortune 1,000 teams across 37+ markets, and builds conversion work on proprietary datasets developed over a decade of campaigns rather than scraped. He writes Bionic Business, read weekly by thousands of subscribers.

More about Sam  ·  LinkedIn  ·  X