How to measure automation honestly when time-saved numbers are usually invented
Why vendor hour-savings fail a controller rebuild test, and how volume times minutes times loaded rate stays honest without fake case studies.
- roi
- measurement
- operations
- finance
If a controller cannot rebuild the figure from the words on the page, delete the figure. That is the whole method. Most "time saved" copy fails it on contact.
Vendors print round numbers because round numbers convert: save 10 hours a week, cut AP in half, give your team their afternoons back. None of those sentences contain volume, task definition, or a loaded rate. They are not measurements. They are adjectives with a digit.
This matters more for AI products than for ordinary SaaS, because the demo is a clean PDF and your pile is not. A number that was true on the vendor's sample set is not a number about your books. Treating it as one is how a pilot gets a victory slide and still never enters the SOP.
Why the usual figures are invented (even when nobody is trying to lie)
Volume is omitted. "Hours per month" without "at N documents" is a marketing number. 60 invoices × 12 minutes is a formula. "Teams save hours on invoicing" is not. If your volume is 20, the hours fall with it. Keeping the 60-invoice figure is the most common self-deception in a buying deck.
The task is not the task. "Invoice processing" might mean open, key, code, route, chase exceptions, and argue with the vendor. A product that drafts the bill and leaves exceptions in a queue did not take the chase. If the card counts chase hours, the vendor must say so. If your team still chases, those hours are still yours.
The baseline is wishful. People remember the painful week, not the median week. They also count interruptions as if a tool removes them. A realistic baseline is a time study on a normal week, or a sample of N items with a stopwatch, or a reconstruction from timestamps in the system of record. Recollection in a workshop is not a baseline.
Loaded cost is replaced with wage, or with a number that feels conservative. Fully loaded cost includes benefits, management, tools, and the fact that the person does not spend 100% of paid time on this task. If you do not know the rate, ask finance. Do not type $25 because it makes the payback look modest. Do not type $200 because it makes the buy look obvious.
Quality and risk are converted into hours they are not. Collections cash is not hours. Do not turn a dunning agent's reminder cadence into "cash collected." Do not turn a lower error rate into "headcount reduction" unless you actually will not backfill. Do not annualize a one-week spike.
AI accuracy is treated as 100%. It is not. Anyone claiming hallucination is solved is lying. Remaining review time is part of the model. If finance products default to human approval — they should, for money movement — you still pay for the click. The honest saving is draft-and-code time, not the whole AP department.
A method that survives a rebuild
Write one row per job. Not per vendor.
1. Name the output artifact. "Draft bill in Bill.com, coded, duplicate flagged." "Reminder sent from the AR mailbox, note on the invoice." "Invoice PDF sent from our domain, same document in QuickBooks." If you cannot name the artifact, you cannot measure the job.
2. Count volume for a representative month. Pull it from the system of record. Do not guess. Note seasonality in a sentence, not a multiplier you cannot defend.
3. Time the current path on a sample. Ten to twenty items is enough to get a median and to see the fat tail. Split the minutes: open, key, code, route, chase. The product will not eat a step you did not list.
4. Hours = volume × minutes / 60, plus named chase or review you still expect. Example using catalog math for Billtray (not a customer study): 80 bills/month × 8 minutes to open, key, code, and route = 640 minutes ≈ 10.7 hours, plus about 4 hours of exception chase ≈ 15 hours at that volume. If you already have capture that keys cleanly, do not buy a second product on top of hours that are already gone. If you process 40 bills, cut the formula in half; do not keep 15.
Second example, Ledgerline: 60 invoices/month × 12 minutes to draft, attach the PO or work ticket, PDF, and email = 12 hours/month at that volume. At 20 invoices, 4 hours. If you send fewer than about 20 a month, a QuickBooks recurring invoice is cheaper.
5. Dollars of named work = hours × your loaded rate. At $70/hour, 12 hours is $840/month of named work at 60 invoices. That $70 is an input you supply, not a number we measured in the wild. There is no third number from "companies like you."
6. Subtract what remains. Approval clicks. Exception queue. The awkward call. Retraining when a vendor's PDF changes. Model-provider changes (see the companion article on deprecation). Setup time in the first month, as a published number, not "you'll be live immediately."
7. Compare to price, including usage. Unattended agents have a cost of goods. Credits or metered runs are the unit. A silent overage bill is a bug in the vendor's design, not a badge of "success." Hard spend ceilings belong in the comparison.
8. Run it on your own data before you treat the formula as true. Catalog math is a hypothesis with volume printed next to it. Your documents are the experiment. If week one does not show the hours at your volume, cancel.
What not to put on the slide
- Headcount you will not actually remove.
- A payback period that assumes a customer result you do not have.
- Aggregate "hours recovered across our customers." We will not print that until operations can defend it from production telemetry.
- Accuracy percentages. No product has a published eval (dataset, sample size, date, named failure mode). Until one exists, do not invent a rate.
- NPS, "loved by," or a count of companies.
How this should appear on a product card
Hours are catalog math at a stated volume, never a trophy. The volume assumption sits next to the hours. The refusal sits next to the claim: it will not pay the bill; it will not sue anyone; it will not invent a late fee the contract does not allow.
If nox.markets ever shows a case study, it will need a named company (with permission), a size band, a named job, and a figure a controller can rebuild. Until that exists, the ROI model is hours × loaded rate minus price. That is enough to decide. It is not a transformation story, and it is not trying to be one.
Use the same sheet for build vs buy. The formula does not care who wrote the code. It cares whether the minutes left your week, in a system you already open, without creating an un-auditable posting.
A dishonest card next to an honest one
Dishonest: "Save 10 hours a week on accounts payable." Missing: volume, steps, who still approves, what happens to exceptions, whether the hours are median or the worst week of the year.
Honest, using catalog math (still not a customer study): "15 hours/month at 80 bills × 8 minutes to open, key, code, and route, plus about 4 hours of exception chase. It will not pay the bill. New vendors sit in exceptions until someone maps them. If you already capture cleanly, do not buy this on top."
The second version can be wrong for you — if your volume is 25, it is wrong in a way you can fix in a minute. The first version is wrong for everyone and cannot be fixed without inventing the missing inputs.
If you are the buyer, refuse to paste the first version into a board deck. If you are the vendor, refuse to ship it. The rebuild test is the same on both sides of the table.
More from the blog
The hidden costs of tool sprawl in mid-size companies
The 15th SaaS login is rarely the fee. It is questionnaires, champions who leave, and no shared context. How 10–200 person companies should gate the next AI SKU.
Data protection questions a buyer should ask any AI product vendor
A usable questionnaire on subprocessors, training, retention, human access, approval, and exit for operators and controllers at 10-200 person firms.
When the model provider retires the model: dependency risk you can actually plan for
Model deprecation is scheduled maintenance. What breaks, who pays to retest, and which contract terms a buyer should require in writing before you buy.
