Why most company AI pilots never reach production
Five delivery failures keep AI pilots out of the SOP. How a 10–200 person company can inspect ownership, integration, and trust before another demo.
- operations
- procurement
- ai-pilots
- delivery
The operations lead at a 10–200 person company has usually been told, since 2024, that AI will take load off the back office. In 2026 the typical stack looks like this: a ChatGPT Team seat, a Copilot licence that nobody opens, one Zapier zap that posts to Slack, and nothing doing work they used to pay a person to do.
That is not a belief problem. She believes. It is a delivery problem. The demo was never the hard part. The hard part is the gap between "this looked good on a slide" and "this is in the SOP, writes to the system of record, and someone can answer for it when it is wrong."
That gap has five nameable failure modes. They are not mysterious. They are the same five reasons a second pilot will fail if you do not inspect them before you start.
1. Pilot purgatory
The pilot demoed well. Then the real questions arrived:
- Who owns it after the vendor's success manager leaves the Slack channel?
- What happens when it is wrong — a wrong GL code, a wrong customer name, a reminder sent to a closed account?
- How does the output reach QuickBooks, HubSpot, the warehouse system, or the shared mailbox the process actually lives in?
- Who retests it when the vendor ships a new model?
If those answers are missing, the thing never crosses from "impressive" to "in the SOP." It sits in a folder named Pilot. The licence fee is the cheap line on the P&L. The expensive line is four months of the operator's attention, plus the political cost of telling the CEO it did not ship.
A useful test before you fund the next one: write the owner, the exception path, the system of record, and the retest owner on one page. If you cannot name them, you are buying a demo, not a process.
2. No in-house ML people, and no path to hiring them
A 90-person company will not hire an ML engineer. The two engineers who could stitch an agent together are the two keeping billing alive. Every AI initiative competes for the scarcest resource in the building.
"Just build it with the API" is advice from people whose day job is building things with the API. The call is the easy 20%. The rest is connectors that break, evals that nobody wrote, failure handling, audit logs, and re-testing when a model changes. That is not a weekend. It is weeks of the people you cannot spare, then a slice of them forever.
If the job is your differentiator — the quoting logic that is the company, the routing that only your warehouses understand — those weeks may be the right spend. If the job is reading a vendor PDF and opening a draft bill, you are staffing a product team for commodity work. The failure mode is not "we cannot code." It is "we staffed the wrong work."
3. Tool sprawl
There are already 14 SaaS subscriptions. Each adds a login, an invoice, a security questionnaire, and a champion who will eventually leave. The 15th tool must be dramatically better than the 14th to clear procurement fatigue. Most AI point products are marginally better at one thing, and still leave the last stretch into the books to a human.
Sprawl is not vanity. It is operational tax: another admin, another DPA, another "who has the password" thread, another tool that does not share context with the other thirteen. A new AI SKU that cannot replace two existing tools, or cannot write into a system you already open, is adding surface area. Surface area is how pilots die quietly: nobody turns them off, and nobody puts them in the SOP either.
Before you add the 15th: list what it replaces, what it writes to, and who owns the questionnaire. If the honest answer is "it is extra," it will not survive Q2.
4. Integration is where the money goes
The model is the cheap part. The expensive part is QuickBooks, HubSpot, the shared Outlook mailbox, eleven years of contracts in Drive, and the Excel file the quoting process depends on.
Generic AI gets a long way on a clean PDF in a demo. It leaves the last stretch — the part that touches systems of record — as a consulting exercise or as six months of internal thrash. Custom remainder work of that kind often lands in a $20k–$60k band, or the calendar equivalent. That is a planning range for skipped integration, not a measured customer result. It is the cost of the work the demo skipped.
The last stretch is the whole job. An invoice that is 90% extracted and 10% unposted is still a person in AP. A ticket that is "drafted" in a chat window and never lands in the helpdesk is still a person on the laptop at 9pm.
Inspect integration before you inspect the model. Ask: which objects does it write, with which scopes, and can you revoke the grant. If the vendor's answer is "export a CSV and someone uploads it," you have not left the pilot.
5. Trust and compliance blockers
She cannot put customer PII into a tool she cannot answer questions about. The insurer asks about subprocessors. The largest customer's MSA has a data-handling clause. The CFO asks what happens if the AI books a wrong invoice amount. "The model is usually right" does not file with an auditor.
This is the failure mode that kills finance and healthcare-adjacent work after the technical demo succeeded. It is also the one vendors paper over with padlock illustrations and words they have not earned. SOC 2 Type II, ISO 27001, and HIPAA are certifications or attestations. If the vendor does not hold them, they must say so. A dated roadmap is allowed. A present-tense "we are certified" is not, unless the report exists.
For money movement, the structural answer is not "trust the model." It is confidence scores, human approval as the default, a citation back to the source document, a named failure mode, and an exportable audit trail. A published error rate on a named eval is what a buyer should demand; we do not have one yet. You approve. You do not trust.
If those artifacts are missing, the pilot will stall in legal even if ops loved the screenshot.
What the five modes share
They are not model failures. They are ownership, staffing, procurement, systems-of-record, and governance failures. Horizontal assistants (ChatGPT, Claude, Copilot) are good at capability and bad at process: nothing runs when nobody is typing, nothing writes to the ERP, and the good prompt lives in one employee's head until they resign. Automation canvases are good at pipes and still leave you as architect, QA, and on-call. Agent builders are workshops. A workshop is a project.
The buyer at this size of company does not have a project office for AI. She has a week, a budget band, and a process that already has a name: invoices, bills, reminders, routing, close.
The inspectable questions, then, are boring on purpose:
- Who owns it in month four?
- Which system of record does it write to, and can you revoke that?
- What does it refuse to do?
- What was it tested on, and what is the named failure mode?
- What happens when the model changes?
- Can you pause it, export the log, and cancel without unwinding work already in your books?
If a vendor cannot answer those in writing, you are not behind on AI. You are being invited back into purgatory.
nox.markets sells finished jobs — tools, agents, workflows, and packs — with integration treated as the product, not as a phase after the pilot. Hours on the card are a time-saved rationale you can check. A published eval is planned. Keep the seats; they make people faster at work they already do. A process that has to run without that person at the keyboard is a different layer. We will not pretend a demo is that layer.
More from the blog
The hidden costs of tool sprawl in mid-size companies
The 15th SaaS login is rarely the fee. It is questionnaires, champions who leave, and no shared context. How 10–200 person companies should gate the next AI SKU.
Data protection questions a buyer should ask any AI product vendor
A usable questionnaire on subprocessors, training, retention, human access, approval, and exit for operators and controllers at 10-200 person firms.
When the model provider retires the model: dependency risk you can actually plan for
Model deprecation is scheduled maintenance. What breaks, who pays to retest, and which contract terms a buyer should require in writing before you buy.
