Why AI pilots stall before production, and the five checks that get them out
The short version. Most AI pilots stall because nobody designed for production: no agreed success metric, data the system can’t reliably reach, no cost per task, no permissions model, and no monitoring. MIT found 95% of generative AI pilots show no measurable P&L impact. Settle those five things before the pilot starts and it has a route to production.
The demo worked. Everyone saw it work. Six months later it still isn’t in production, the team has moved on to the next pilot, and the board is asking what the AI budget returned. If that sounds familiar, you have plenty of company. MIT’s NANDA initiative studied 300 public AI deployments and interviewed 150 leaders, and concluded that about 95% of generative AI pilots deliver no measurable impact on the P&L.12
Agents won’t fix this on their own. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.3 Not one of those three is a model problem. All three are engineering and management problems, which means they can be designed out before the pilot starts.
The model is rarely the problem
When a pilot stalls, the instinct is to blame the model and try a newer one. It rarely helps. MIT’s researchers found the core barrier wasn’t model quality, infrastructure or regulation. It was that most systems don’t retain feedback, adapt to context or improve over time. They also found that externally partnered projects succeeded about twice as often as internal builds.2 In practice, a pilot is usually built to prove the model can do the task once. Production needs a system that does the task every day, at a known cost, within known limits, and gets better.
That gap is predictable, so it can be checked. Before we write code on an engagement, we work through five questions with the people who will own the result. If any answer is “we’ll figure it out later,” the pilot is a demo, not a route to production.
1. Is there one number that defines success?
Not “improve customer service.” A number with a baseline: tickets resolved without a human, hours of manual review removed per week, days cut from month-end close. Agree it with the person who signs the budget, before the build, and measure the pilot against it on real data. If the number doesn’t move, you stop, and you’ve spent weeks rather than quarters finding out.
This one check also answers the CFO’s question in advance. A pilot without an agreed number can only ever be judged on impressions, and impressions don’t survive a budget review.
2. Can the system reach the data it needs, reliably?
Pilots often run on an exported spreadsheet or a hand-cleaned sample. Production runs on live systems, with their missing fields, changed schemas and permission rules. Gartner predicts that through 2026, organizations will abandon 60% of AI projects that aren’t supported by AI-ready data, and found that 63% of organizations either lack, or aren’t sure they have, the right data management practices for AI.4
The check is concrete: name every system the AI must read from or write to, confirm there’s an API or pipeline it can use in production, and confirm who owns each source. If the answer for a critical source is “someone exports it weekly,” fixing that is the first deliverable, not a phase-two item.
3. What does one task cost, and who is watching it?
Model calls, retrieval, tool use and retries all cost money, and agentic workflows multiply them. A pilot that costs a few dollars a day in testing can cost thousands at production volume. Gartner names escalating costs as one of the three reasons it expects agentic projects to be cancelled.3
Measure cost per completed task during the pilot, next to the success rate. Then you can compare it with the cost of the work as it’s done today. If the AI version is more expensive per task and not meaningfully better, that’s an answer too, and a cheaper one than finding out after launch.
4. What is the system allowed to do, and where does a human approve?
An assistant that drafts an answer is low risk. An agent that issues a refund, changes a record or emails a customer is not. Security and risk teams stall projects, rightly, when nobody can say what the AI can touch. Design it up front:
- Give the agent its own identity, not a shared admin key.
- Register every tool it can call, with scoped permissions.
- Put human approval exactly where an error is expensive, and nowhere else.
- Log every action so any decision can be traced afterwards.
Done this way, governance speeds approval up. The risk conversation happens once, against a concrete design, instead of in every steering meeting.
5. How will you know when it gets worse?
A model that worked in March can be wrong in June. Input data drifts, providers update hosted models, and a small prompt change breaks an edge case. Without evaluation and monitoring, the first sign is a complaint. Production needs a test set the system must pass before every release, live tracking of accuracy, latency and cost, and a named owner who gets the alert. This is also how a system “retains feedback and improves over time,” the capability MIT found most deployments were missing.2
What to do with the pilot you already have
Run the five checks against it. Most stalled pilots fail two or three, and they’re usually fixable: agree the number, build the missing data access, add cost tracking, design the permissions and stand up evaluation. Sometimes the honest result is that the use case doesn’t pay at production cost, and the right move is to stop and pick a better one. Both answers beat another quarter in pilot purgatory.
The pattern behind every check is the same. Treat production as the deliverable from the first week, not the step after the demo. The model is a commodity. The system around it is the work.