The pattern is consistent enough to be predictable. Someone builds a proof of concept, it does something genuinely impressive in a meeting, everyone agrees it is the future, and four months later it is not running. No decision was taken to stop it. It just quietly stopped being anyone's job.
The usual explanation is that the technology was not ready. That is rarely what happened. What happened is that the thing being tested was never the thing that had to work.
A demo is a performance with a friendly audience
Whoever runs the demo chose the example. They know which inputs behave, they are watching while it runs, and if it wobbles they can adjust and try again. Every one of those conditions is absent at four o'clock on a Tuesday when a real enquiry arrives in a format nobody anticipated.
So a demo answers a question that was not really in doubt. Can the model do this kind of task? Almost always yes. The open question is different and much harder: will this job get done, correctly, every time, when nobody is watching and nobody remembers it exists.
Five things a demo never has
The distance between the two is mostly plumbing, and it is the same five pieces almost every time.
None of this is the interesting part of the work, which is exactly why it gets skipped. It is also the entire difference between a capability and a system.
The pilot that belongs to nobody
There is a second failure underneath the technical one, and it is the more common of the two.
Pilots usually get run as experiments, by whoever was curious enough to start. That works while their attention holds. Then a quarter gets busy, the experiment goes unattended, and nothing anywhere in the business registers the loss, because nothing in the business had come to depend on it. The pilot was never load bearing. It was running alongside the real process rather than inside it, which meant the old way had to keep running too, and everyone kept using the old way because it was the one that was definitely going to work.
Running a new system in parallel with the old one feels like the careful choice. It is often what guarantees the new one never gets adopted: as long as the old path still works, there is no week where anyone has to rely on the new one.
This is also why success criteria matter more than they sound like they should. Agreed in advance, in writing, they turn "is it working" from an opinion into a check. Without them, a system that is quietly doing its job gets cancelled by whoever is looking at the invoice, and a system that is quietly doing nothing survives because it is not annoying anyone.
The test that settles it
If this stopped running for a week, would anyone notice before you told them?
If the answer is no, it is not in production. It is a demo on a timer, and the effort is going into something that will never show up in the numbers.
If the answer is yes, and you can name who notices and what they would be missing, the pilot has already crossed the line that most never do. Everything after that is refinement, and refinement is cheap.
Which is the argument for starting smaller than feels ambitious
One job, chosen because it repeats and because a missed one costs money. Give it the five pieces above and let it run for a month against a number you wrote down at the start. That is a boring proposal next to a platform that does everything, and it is the version that is still running in a year.
The businesses getting real value out of this are not the ones with the cleverest tools. They are the ones where something unglamorous has been running every day for six months and has a record to show for it.