The Plumbing Problem Why pilots succeed and production fails, and why the failure is almost never the model. A structural account of five failures that appear, in roughly the same order, in almost every enterprise AI deployment we have been asked to rescue. With what each one costs and what it takes to fix. Radexus. Institutional memory for revenue. radexus.com This paper argues a single claim: the reason enterprise AI pilots convert to production at such poor rates has almost nothing to do with model capability, and almost everything to do with the state of the data substrate underneath them. We call that substrate the plumbing, and it fails in five identifiable ways. ## Method The observations here come from discovery engagements across manufacturing, distribution, insurance distribution, logistics and healthcare provision, in India and the United States, between 2024 and 2026. In each case we had read access to systems of record and spent two weeks establishing what could and could not be computed from data the company already held. We are describing what we found, not surveying a market, and the sample is biased towards companies that had already tried something and been disappointed. ## Failure one: the same thing has several names A customer exists in the ERP as a legal entity, in the CRM as a spelling, in the dealer portal as an abbreviation and in the service application as a site address. Every report built on top is quietly wrong, and every person defending their own report is quietly right. The consequence is not merely inaccurate reporting. It is that no coherent history of any customer exists, which means no pattern can be formed across their episodes, which means the system cannot learn anything about them. Resolution is not a data hygiene task. It is the precondition for memory. ## Failure two: nobody wrote the definitions down Open pipeline, stalled quote, dormant account, active dealer. Everyone uses the words. No two departments compute them identically. A system deployed on top of this does not resolve the disagreement, it industrialises it and gives both sides faster ammunition. In our engagements the definition work reliably consumes the first week and reliably returns more than anything else in the fortnight. The exclusions are where the value sits, because the exclusions are where the disagreements actually live. ## Failure three: there is no lineage A number appears on a screen. Nobody can say which system it came from, when it was read, or what transformed it. The first time it looks odd to someone senior, it stops being trusted, and trust does not return once lost. We treat traceability to source within one minute as a hard requirement rather than a feature. ## Failure four: nothing is allowed to write back Output lands in an export that three people open for a month. Acting on it means retyping into the system where work happens, so nothing is acted on. The refusal to let software write into a system of record is entirely rational and the answer is not to argue against it, but to build write-back with named approvers, dry runs, audit trails and rollback windows. ## Failure five: nothing is measured after go-live No thresholds, no evaluation suites, no gates. Quality drifts because the world moved rather than because the code changed, and the first person to notice is someone senior losing confidence in a number. We have not encountered a failed deployment that had evaluation gates with owners attached. [The compounding claim] Fixing the plumbing is expensive once and cheap thereafter. In our engagements the second play typically costs about a third of the first, because the resolution, definitions and lineage the first one paid for are already there. This is the entire economic argument for doing it properly rather than shipping a demo. ## What follows from this If the analysis holds, then the correct sequencing for any enterprise deployment is: agree a number, resolve the entities, settle the definitions, ship one play against the number, gate the release, and only then consider the second. The industry has largely inverted this, starting with a capability and looking for somewhere to apply it, which is why so much of it stalls at the pilot. (c) 2026 Radexus, a Global AI Forum company. San Francisco and Chennai.