The short answer
Pilots stall because production is a different object, not a bigger demo: it adds integration, exceptions, evals, access control, and an owner. The crossing is six deliberate steps: shrink to one owned workflow, define production on day one, build inside the workflow, put evals in before launch, ship to a small set of real users in weeks, and operate with a full handover. Skipping steps is how the 88 percent happens.
Every stalled AI pilot was once a good demo. The difference between the systems that cross and the ones that decay in a status deck is not enthusiasm or budget. It is that somebody scoped the crossing itself.
Why pilots get stuck
Gartner's prediction pointed the same direction with the causes attached: at least 30 percent of generative AI projects abandoned after proof of concept, on poor data quality, inadequate risk controls, escalating costs, or unclear value. The full failure anatomy is in why AI pilots fail; the short version is that a pilot proves the model can do the task, while production requires everything around the task: the integrations, the exceptions, the proof it keeps working, and the person accountable when it does not. Every number in this genre, with what each study actually measured, is reconciled in the failure statistics reference.
A pilot proves the model can do the task. Production is everything around the task.
The reframe that unlocks the crossing
Stop treating production as the pilot's reward. Treat it as the specification. The question is never "did the demo impress" but "what exactly is missing between this and a system real users depend on": name those items, and the crossing becomes a build plan with an end date instead of a hope.
The six-step crossing
- 01Shrink to one workflow with an owner. One queue, one process, one person who can say yes. Volume enough to matter, data reachable. Pilots that serve committees serve nobody.
- 02Define production on day one. Named users, the systems it must integrate with, the exceptions it must survive, the eval thresholds it must pass, who carries the pager. Write it down before building; this document is the scope.
- 03Build inside the workflow, not beside it. Design against the observed process, with the people who run it, on real data from the first week. Systems built from briefs automate an imagined workflow and stall on contact with the real one.
- 04Put evals in before launch, not after. A golden set of real cases, boundary tests for what the system must never do, thresholds that gate release. Without them, production readiness is a feeling.
- 05Ship to a few real users in weeks. A first production cohort beats a bigger demo. Real usage surfaces the exceptions no workshop predicts, while the blast radius is still small.
- 06Operate, then hand over completely. Monitoring on silence as well as errors, runbooks, model-lifecycle plans, and a transfer of code, prompts, infrastructure, and evals into your accounts. Production that only the vendor can run is a pilot with billing.
Steps four and six have their own field guides: evals for business buyers and the ownership handover. Both exist because they are the two steps vendors most reliably skip.
Doing this at operations scale
The studies above measure enterprises, where the crossing fights procurement cycles and committee ownership. At operations scale the structural advantages flip your way: the person who runs the queue is reachable, approval is one conversation, and the whole six-step crossing fits inside a fixed-scope engagement measured in weeks. That is precisely the shape of an embedded engagement: discovery inside the workflow, a production pilot with real users, and the handover at the end. The 88 percent is not destiny. It is what happens when nobody scopes the crossing.