One of my clients, a SaaS company doing about $4.2M a year, had an upgrade agent with great numbers. Over three months it drafted 212 upgrade conversations, and 31 of them closed, adding $214K in annual revenue.
It also declined to act roughly 338 times, and nobody knew. When the agent decided not to act, it wrote nothing to the log. The run came back empty and the workflow ended.
The main cause was boring: a seat-count field left blank on accounts set up before a 2025 change to the database. 94 accounts that should have heard from the company never did. Eleven of them later downgraded and six cancelled. Every review in those three months looked only at the cases where the agent acted, so every review said it was doing well. The 338 had to be rebuilt afterwards by replaying the runs, so treat it as approximate.
An automation that breaks usually breaks loudly: a failed step, an error email. An agent that’s wrong, or quietly doing nothing, looks the same as one that’s working. That’s why an agent needs a short weekly routine and a handful of numbers, and why the first rule is to log what it didn’t do.
1. Why this matters now
Two things changed.
Agents are cheap enough to put on real work. Small businesses now run them on enquiries, upgrades, support replies and proposals that go out to real people.
And the model underneath can change without you touching anything. In 2023, researchers at Stanford and Berkeley ran the same tasks through the March and June versions of GPT-4 and GPT-3.5 and found that “the behavior of the ‘same’ LLM service can change substantially” within a few months (Chen, Zaharia and Zou). In April 2025, OpenAI rolled back an update to GPT-4o in ChatGPT within a week of release because it had become “overly flattering or agreeable”.
You can pin a model version in most APIs, which helps. But versions are retired, and eventually you’ll want the newer one. The prompt you tested in March may be running on a different model by autumn.
Most advice treats measurement as something you do at launch: test it, ship it, move on. Several of the costly failures I’ve documented came after launch, from data that quietly broke, like the one above. An agent that runs on its own doesn’t watch itself. Measuring has to be a weekly habit, small enough that you keep doing it.
2. The measurement routine, step by step
Step 1: Write down the job, what “correct” means, and what it’s for
One sentence for the job: “Draft an upgrade conversation for every account that crosses its tier’s break-even point.” Then a few sentences on what a correct output looks like, specific enough that two people would grade the same run the same way.
Then one more line: the business result this job exists to move. Revenue, signed clients, retained customers. You’ll need it for the seventh metric, and leaving it out is behind one of the failures below.
Step 2: Build a test set
Collect 20 to 50 past cases where you know the right answer. Include the easy ones, the hard ones and the strange ones. Label each with the correct output.
A $1.1M Shopify app business I work with did this before launching a support agent. They ran it on 50 messages drawn from real support history. Ten of the replies would have caused trouble. Six mentioned a pricing tier the company had retired months earlier, because the feature file it read hadn’t been updated since. Four promised features that engineering had never agreed to build, things like “that’s coming next month”. No customer saw any of it. The test took one afternoon, and the founder now runs one before any agent goes live.
A test set built from past messages can miss what customers will ask once the agent is live. That’s why it’s only the start.
Step 3: Log every run, including the ones where it did nothing
For each run, save: the input, the outcome, the output, the prompt version, the model, the cost, the time it took, and whether a person changed the output afterwards.
The outcome field is the one that would have saved the SaaS company above. Give every run one of three outcomes: it acted, it chose not to act and gave a reason from a fixed list, or the case wasn’t eligible. A run that produces no record is a failure in itself.
Another client, a micro-SaaS with about $980K in annual revenue, built the same kind of upgrade agent this way from the start. In its first ten weeks it drafted 88 conversations and declined 151. Each Monday the founder spent about fifteen minutes reading the declines, grouped by reason. 64 of them shared one reason: a usage field that stopped filling in for customers who connected through the API instead of the dashboard. That fault had existed for over a year. They found it in six weeks. Once it was fixed, those 64 accounts produced 19 conversations and 7 upgrades, worth $38,400 a year. That figure assumes those seven wouldn’t have upgraded anyway, which is likely but unproven, since the founder knew three of them well.
Same kind of agent, same job, two businesses. One logged its silence and found a year-old fault. The other didn’t, and lost three months.
A spreadsheet row per run is enough to start. Most automation platforms can append a row as the last step.
Step 4: Review a sample every week
For the first two weeks of any new agent, review every output. Once the output has been clean for six to eight weeks, move to a sample audit. One version that works: review everything one workflow produced on a Friday, alternating workflows each week. A random pull of ten runs a week works too. Either way, also check anything the agent flagged as low confidence, and keep a running log of any mistake that reaches a customer.
What matters is that the sample isn’t chosen by the problems. If you only look at runs that were flagged or that someone complained about, you’re measuring complaints.
And approving isn’t reviewing. A $1.9M online education business ran an agent on pre-purchase questions during a launch, with a person approving every reply. Of 1,104 drafts, 1,081 were approved, 97.9%. The median approval took 7 seconds, and 412 took under 4. Three approved replies described guarantee terms that didn’t exist, including a 90-day full refund against a real 14-day conditional guarantee. All three people bought. It cost $2,994 in terms the business chose to honour, plus a lost chargeback of $1,497 because the buyer had the email. The reviewer was a part-time VA with no authority over refunds and no written copy of the terms.
A 7-second median doesn’t prove the other 1,078 replies were wrong. It proves nobody was really checking. The business has decided to end one-click approval for any reply that touches a guarantee, refund, payment plan, deadline or bonus. The plan is to route those to the person who owns the offer, and keep one click for everything else.
Step 5: Re-run the test set after any change
Changed the prompt? Re-run it. Switched models? Re-run it. Provider announced an update? Re-run it. Compare the score with last time before the change goes live.
Step 6: Keep one simple dashboard
One row per week, one column per metric below. A sheet works. The point is to see the trend, so you notice when a number moves.
3. The seven metrics
Task success rate
The share of runs that finished with a correct, usable output.
Task success rate = correct runs in the sample ÷ runs in the sample.
This comes from your weekly review. The automation platform counts runs that completed. You’re counting runs that were right.
Accuracy against the test set
The share of test cases where the agent’s output matches the labelled answer.
Accuracy = matching test cases ÷ total test cases.
This is your before-and-after number for every change. It’s only as good as the test set, so add new cases whenever the weekly review finds something the set doesn’t cover.
Override rate
The share of reviewed runs where a person changed or rejected the agent’s decision.
Override rate = runs a person changed ÷ runs a person reviewed.
This is the closest thing to a live accuracy score you get without grading everything yourself. The consultant in my lead qualification workflow has overridden her agent 7 times in 83 enquiries, about 8%. More on what to do with that number in section 6.
Abstention and escalation rate
The share of runs where the agent chose not to act, or handed the case to a person.
Abstention rate = runs it declined ÷ eligible runs. Escalation rate = runs passed to a person ÷ total runs.
Watch both directions. Too high, and the agent isn’t saving much time, or something upstream is broken, as the micro-SaaS found. Too low, and it may be guessing on cases it should have passed on. A sudden drop after a prompt change is a warning sign. The micro-SaaS wrote its thresholds down: the agent takes on new work only once declines stay under 25% of eligible accounts for four weeks in a row, and any week above 40% triggers a data check before anything else.
Errors that reach customers
The count of mistakes someone outside the business saw: a wrong reply, a wrong price, a promise you can’t keep. Log each one with the run it came from.
This is the slowest metric to move and the one that matters most. One of these deserves more attention than a two-point drop in accuracy. When one happens, put that workflow back on full review until you know why.
Cost per successful run
What the agent costs, spread over only the runs that were right.
Cost per successful run = total spend for the period ÷ (total runs × task success rate).
Cost per run on its own flatters a bad agent. If half the runs need fixing, the real cost of each useful output is double what the dashboard shows. For the full cost picture, see how much an AI agent costs.
The business result
The number the job exists to move, measured on its own schedule.
This is the one people skip, because it’s slow. A YouTube-first media business doing about $740K a year gave an agent the job of choosing video titles and thumbnail text, and judged it on predicted click-through. Over 14 uploads, click-through rose from 5.8% to 7.6%. View duration on videos of comparable length fell from 7 minutes 12 seconds to 4 minutes 48 seconds, and affiliate revenue per thousand views fell 31%. Two sponsors raised weak post-roll results at renewal, and one cut its commitment from four slots to two. That sponsor gave budget as the reason, so the link to the titles is a reasonable guess rather than a proven one.
The agent did exactly what it was measured on. The fix they’ve settled on: it only proposes titles, a person picks, and the plan is to judge it on affiliate revenue 28 days after upload instead of click-through.
4. What it replaces
Spot checks when you remember, and the feeling that it’s probably fine.
5. What it needs
Run logs with an outcome on every run, including the ones where the agent did nothing. A labelled test set. A column for “a person changed this”. And about fifteen minutes a week to start, which is what the micro-SaaS founder spends on Mondays.
6. The number to watch
The override rate, with the abstention rate next to it.
Overrides come from real work at full volume without any extra grading. Every time you edit or reverse the agent’s decision before it goes out, that’s a data point. Abstentions tell you what the overrides can’t: the cases where it never produced anything for you to override.
Watch the trend more than the level. An override rate that falls as you fix the prompt and then holds steady is healthy. One that creeps up week after week means the inputs or the model have shifted, and it can move before customers notice.
Use it to decide what the agent is allowed to do alone. The consultant’s agent can send declines without her only after 40 decisions in a row with no override. The counter has reset twice so far.
To translate these numbers into money, the AI agent ROI post covers baselines and the dashboard for that side.
7. Where it breaks
Measuring only what the agent did. The SaaS company at the top had three months of healthy reviews because every review left out the runs where the agent did nothing.
Sampling the easy cases. If the sample is only what got flagged, or always the quietest slot of the week, you’re grading the easy work. Rotate what you look at.
Metrics the agent can be pushed to game. Tell a prompt to escalate less and it can, by guessing more. Optimise on click-through and you get clicks. Read every metric next to the business result, never on its own.
Approvals nobody reads. A 98% approval rate at seven seconds a draft is a rubber stamp. Send the drafts that can cost money to someone with the authority and the facts to judge them.
Overrides nobody records. If a person fixes the output in the email client instead of in the workflow, the override never reaches your log, and the agent looks better than it is. Make the fix happen where the log can see it.
Model updates. The reason the test set exists. Re-run it on any provider announcement, and pin versions where your API allows it.
Before this week is out, add an outcome column to whatever your agent logs, with “did nothing, and why” as one of the options. It’s the cheapest metric to start and the one that finds the problems nobody is looking for.
Related
- AI agent ROI for the money side.
- My AI lead qualification workflow shows the override rate on a working agent.
- How to turn an SOP into an AI agent covers testing a build against past cases and in shadow before it goes live.
- AI agents for operations and workflow for testing and cost controls in an operations setup.
