The Demo Was Never the Hard Part: Budgeting for the PoC-to-Production Gap
A working AI prototype came together in two weeks. It took another nine months to reach real customers, and nobody had planned for them. What actually eats the timeline between a demo that works and a system people can use.
The demo is the easy 10%. Everyone budgets for the demo.
A few years ago I watched a working AI prototype come together in about two weeks. Everyone in the room was happy. Screens shared, nods all around, a couple of people already talking about the launch date like it was a formality.
That prototype took another nine months to reach real customers. The two weeks were the fun part. The nine months were the actual job, and nobody had planned for them.
You will see a figure passed around that something like 90% of AI projects never make it past the proof-of-concept stage. I have never been able to trace that one to a source I trust, so treat it as folklore rather than data. What I can tell you is the shape of the thing from the inside, because I have shipped a few of these and watched more get quietly shelved.
The demo lies to you
When I worked on embedding an insurance product into Microsoft Copilot, the working prototype was maybe 20% of the timeline. The rest went into data contracts, evaluation sets, failure handling, and a long stretch of convincing everyone in the building that the model would not say something strange to a paying customer on a bad day.
None of that shows up in a demo. A demo runs on three clean inputs you picked yourself. Production runs on whatever a real person types at 2am, plus the input someone fat-fingers, plus the edge case nobody imagined.
Where projects quietly die
Problem definition is short on paper. Two to four weeks, usually. It is also where most doomed projects get their fatal flaw. If your success metric is vague, everything downstream inherits the vagueness. “Improve customer experience” is not a metric. “Cut first-response time on claims by 30% without raising escalations” is something you can build against and test.
Picking the wrong problem well is worse than picking the right problem badly. You can fix a rough build. You cannot fix a project that should never have started.
Then comes data collection and preparation, and this is the phase everyone underestimates by a mile. In my experience it eats most of the total effort, somewhere in the 60 to 80% range, and any estimate you write down for it is optimistic the moment the data touches anything regulated or messy.
This is the unglamorous middle. Cleaning, labelling, chasing down who actually owns a field in some upstream system, figuring out why 4% of your rows have a timestamp from 1970. Nobody puts this part in the pitch deck. It is most of the work.
The gap nobody puts on the roadmap
Deployment takes about as long as building the model did. That single fact is most of the reason so many pilots die.
Teams plan a timeline that ends when the model is trained and accurate. Then they hit deployment and discover a second full project sitting in front of them, one they never scoped. The budget runs out, the sponsor loses patience, and the successful pilot gets quietly shelved.
The model was never the problem. The plan was.
AI work also runs longer than traditional software, and not by a little. If you scope an AI feature the way you would scope a normal one, you are probably wrong by a factor of two before anyone writes a line of code.
Production finds the failures your tests cannot
Last month I got a fresh reminder that crossing the gap is not the end of it.
On my own bootstrapped product, a model failed in production on the busiest day of a campaign. Every test we had written passed. Every one. Production still found a failure mode I had not thought of, on the exact day I could least afford it.
What saved me was not better testing. It was the monitoring and recovery layer I had built for precisely this moment. The system caught the anomaly, switched to a fallback model on the fly, and sent apology emails with credits to the handful of users who got hit before the switch. The users barely noticed. I noticed a lot.
Monitoring exists to catch the silent failures, the ones where accuracy quietly decays over weeks and no error ever fires. The model does not crash. It just slowly gets worse while everyone assumes it is fine. By the time a human notices, you have been making bad predictions for a month.
Three things that move the needle
Strategy over execution. The best engineering cannot save a badly chosen problem. Spend real time here before anyone opens a notebook.
Standardise the stack. Tools like MLflow for experiment tracking and Kubeflow for orchestration exist so that “it worked on my machine” does not become your deployment strategy. Deployment usually takes as long as model development because the path to production was hand-built each time. A standard stack is how you shrink that gap on the next project instead of paying full price again.
Classify your regulatory risk early. The EU AI Act sorts systems by risk level, from unacceptable, which is prohibited outright, down to minimal. Finding out which bucket you sit in during deployment instead of during problem definition can cost you months, or a whole build. For anything in finance, health, or hiring, you want this answered in week one.
What I ask now
These days when someone walks me through a slick AI demo, I do not ask about the model. The model is clearly fine, that is why they are showing me a demo.
I ask what the evaluation set looks like. I ask what happens when the model is wrong, because it will be. I ask who gets paged when accuracy drops and there is no error in the logs. The answers tell me whether I am looking at a product or a science project.
The piece I still have not solved is getting teams to budget for the gap before they fall into it. Everyone plans for the model. The other 80% shows up as a surprise every single time, on every team I have watched, including my own. If you have found a way to make that gap visible on a roadmap before it bites, I want to hear how.
I teach this gap as a seven-day programme now, because reading about it does not close it. Ship in 7 walks one idea from a roadmap to a deployed, hardened application.