Why AI fails in production: the math behind compound error amplification

By
Gabriel De Dominicis

AI fails in production because errors multiply across steps. Five steps at 90% accuracy yields 59% reliability. Here is the compound error math.

Compound error amplification is the phenomenon where accuracy losses at each step of a multi-step process multiply together rather than add, causing end-to-end reliability to collapse far faster than any single-step error rate would suggest. This is the primary reason AI fails in production, not the model, not the data, not the team. The math.

If your model is 90% accurate and your process has five dependent steps, your end-to-end reliability is not 90%. It is 0.9 × 0.9 × 0.9 × 0.9 × 0.9 = 59%. Roughly as good as a coin flip. At ten steps, it drops to 35%. This is why Gabriel De Dominicis, Kapto's MD & Head of AI, frames it bluntly: "If you don't have enough precision in your AI model, you can't even think to automate. Zero automation."

Why AI fails in production: the compound error degradation table

The table below shows end-to-end process reliability as a function of per-step accuracy and number of sequential steps. This is the math that most AI vendor conversations skip.

Per-step accuracy 2 steps 3 steps 5 steps 8 steps 10 steps
70% 49% 34% 17% 6% 3%
80% 64% 51% 33% 17% 11%
85% 72% 61% 44% 27% 20%
90% 81% 73% 59% 43% 35%
95% 90% 86% 77% 66% 60%
98% 96% 94% 90% 85% 82%
99% 98% 97% 95% 92% 90%
99.5% 99% 98.5% 97.5% 96% 95%

This table assumes independent errors per step. In practice, systematic errors, a model that always misreads the same field, are additive on top of this math and can be worse.

Read down the 90% column. Five steps: 59% reliability. That is not an AI problem you can patch. That is a structural consequence of using an insufficiently precise model in a process that requires sequential decisions.

Read across the 95% row. Five steps: 77%. Still not automation-grade for enterprise operations. You would need a human reviewing roughly one in four outputs. At 98%, five steps gives you 90%. At 99.5%, it gives you 97.5%.

This is why Paolo Ferrari, Kapto's co-founder & CCO, puts the automation threshold at 95% as an absolute floor: "The generative AI is not able to give you 95% of accuracy. We at Kapto, yes, we are at 98 end-to-end." The implication scales with complexity: the more entities a document contains, the higher the per-entity accuracy must be for the extraction to come out clean.

The threshold is not a preference. It is the point below which the math makes autonomous operation impossible.

Hand pointing at an analytics dashboard on a tablet showing performance metrics

Three dominant failure modes when AI fails in production

1. The accuracy illusion: demo performance does not transfer to production

The most common deployment pattern: a vendor demonstrates an AI model on curated documents, reports 85-90% accuracy, and the enterprise deploys. In production, the model encounters the full distribution of real documents, edge cases, inconsistent formats, missing fields, handwriting, multi-language content. Accuracy drops. And because the process has multiple steps, the drop is not proportional.

A vendor that one of Kapto's customers evaluated used a large language model for document processing. After eight months in production: 70% accuracy. Per the compound error table, 70% accuracy at five steps yields 17% end-to-end reliability. The automation rate was effectively zero. The same documents were being reviewed manually as before, except now there was a layer of AI overhead on top.

Gabriel's assessment of why LLMs fail here is direct: "70% accuracy of a large language model implies zero automation. Second, a large language model is not learning, so the same problem is there after one moment, two moments. Third, not a monotonic answer. You take the same document two times and you have two different answers." Three distinct failure modes. The first means you cannot automate. The second means the problem does not self-correct over time. The third means an automation system built on top of it cannot be trusted in a sequential workflow, as each step reintroduces uncertainty rather than resolving it.

2. The near-threshold trap: fixing errors becomes more expensive than not automating

The table above has a region that is worse than it looks: the 90-94% accuracy range with four or more steps. In this zone, the system processes enough volume autonomously that the error burden lands on the downstream reviewer rather than being caught upstream. The reviewer is no longer doing their original job, they are auditing AI output, and doing it under worse conditions than if they had handled the work directly.

As Gabriel puts it: "Near the threshold, spotting and correcting errors as an end user is worse than doing everything manually." The errors do not announce themselves. They arrive embedded in otherwise correct-looking outputs, which means the reviewer cannot skim, they must read everything, which defeats the purpose of automation.

This is an accuracy problem and not a workflow design problem. The only fix is higher per-step precision.

3. The entity count effect: complex documents require proportionally higher accuracy

The compound error problem becomes structurally worse as document complexity increases. A simple form with four fields has four entities to extract correctly. An insurance claims document might have forty. A broker reconciliation statement might have four hundred line items, each requiring accurate extraction and cross-referencing against a portfolio.

Each entity is a potential error point. Document processing errors compound with document complexity; a forty-entity claim introduces forty independent error chances. Even a model with 99% per-entity accuracy will produce a fully correct extraction only 67% of the time (0.99^40 = 0.67). At 98% per entity, the probability of an error-free extraction drops to 45%.

This is why the accuracy requirement scales with the complexity of the document type, not just the number of process steps. It is also why "minimum viable models", purpose-built for specific document types and constrained entity sets, outperform general-purpose LLMs on enterprise automation. A general model applied broadly is, as Gabriel describes it, "a drill used as a screwdriver." The tool is not defective. It is being used for a task it was not built to do.

What percentage of AI projects fail in production?

The commonly cited project-failure figures capture the rate of abandonment but not the mechanism. A project can reach "production" and still fail operationally if its accuracy is below the automation threshold. The deployment is live; the automation rate is near zero; the team is reviewing AI output manually.

The correct question is not "what percentage of AI projects fail?" but "what percentage of AI projects achieve the accuracy threshold required for their specific process?" Based on Kapto's observations across European insurance and manufacturing deployments, multiple enterprise AI implementations using frontier LLMs have produced accuracy rates in the 70-80% range after six to twelve months of production operation. Per the compound error table, these accuracy levels make autonomous process execution mathematically impossible regardless of other factors.

The threshold question, "what accuracy do you actually have in production?", is the one question that separates real deployments from proof-of-concept demonstrations. Gabriel's framing for any AI vendor evaluation: "What's the precision of your model right now?" It is, deliberately, a question most vendors cannot answer with a specific number for a specific use case. Anyone who cannot has not deployed their model in a real production workflow.

Frequently asked questions

What accuracy do AI workflows need?

The minimum threshold for autonomous process automation in an enterprise context is 95% per step, and this is a floor, not a target. At 95% per-step accuracy with five steps, end-to-end reliability is 77%, which still requires human review of roughly one in four outputs. For full straight-through automation, per-step accuracy needs to reach 98-99%, depending on process complexity.

The threshold varies by document type and entity count. A process that extracts five fields from a structured form operates under different requirements than a process that extracts forty entities from an unstructured insurance claim and cross-references them against a policy system. The compound error table above assumes independent steps; in reality, correlated errors (the same field consistently misread across documents) can be worse.

The practical implication: "What's the precision of your model right now?" is the correct first question to ask any AI vendor. Anyone who cannot answer it with a specific number for their specific use case has not deployed their model in a real production workflow. Demo accuracy and production accuracy are not the same number.

Why does 95% accuracy fail at scale in multi-step automation?

Because errors compound multiplicatively, not additively. A 95% per-step accuracy with ten steps yields 60% end-to-end reliability which is close to a coin flip. At that reliability level, every other output requires human review, and the cost of that review often exceeds the cost of the process before automation.

The counterintuitive implication is that small accuracy improvements near the threshold have disproportionate impact. Moving from 90% to 95% per-step accuracy on a five-step process increases end-to-end reliability from 59% to 77%, a 30% absolute improvement in usable output. Moving from 95% to 99% pushes it from 77% to 95%, the difference between a human-augmented AI tool and a fully autonomous workflow.

This is why Gabriel frames it directly: "1% of accuracy dropping is killing your automation." It is not hyperbole. At the operating margins of enterprise process automation, a 1% per-step accuracy change moves the system from viable to non-viable.

What is compound error amplification in AI?

Compound error amplification is what happens when AI accuracy losses at each step of a sequential process multiply together. In a process with n dependent steps, each running at accuracy rate p, the end-to-end reliability is p^n. The errors do not add, but multiply. One error plus one error is not two errors; it is one error making the next step's input wrong, which means the next step's error applies to a larger fraction of already-degraded data.

The practical consequence: enterprise domain processes, claims processing, broker reconciliation, order management, invoice matching, are multi-step by design. They were built with sequential validation checks precisely because each step catches errors from the previous one. When AI replaces those steps without the accuracy required to make each autonomous, it removes the error-catching mechanism while introducing new error sources. The result is not incremental degradation. It is collapse.

The threshold argument, and why it sets the floor for AI you can actually deploy

Across Kapto's production deployments, the system has processed over 10,000,000 pages at 98%+ accuracy. The errors that do occur are concentrated in cases where the source document is genuinely ambiguous, cases a human reviewer without additional context would also have misclassified, rather than in routine extraction.

For insurance deployments, the straight-through automation rate reaches 70% of claims handled automatically, with over 95% of communications routed without human intervention. The majority of cases run straight through; supervisors handle the exceptions.

That outcome is a consequence of the architecture. Specifically, the decision to treat 95% accuracy as a floor rather than a target, and to build confidence-calibrated uncertainty into every extraction step so that low-confidence outputs route to human review rather than propagate through the process chain. The system does not attempt to automate what it cannot do reliably, a design principle explained in depth in the confidence engine approach.

This is the operational argument for the 95% threshold. Below it, automation does not reduce the human workload. It redistributes it to harder, more error-prone review tasks. Above it, and particularly above 98%, the compound error table starts working in your favor rather than against you.

Why agentic AI amplifies this problem

The current enterprise AI narrative is agentic: autonomous AI agents that chain together to complete multi-step tasks without human intervention. Gartner projects that 40% of AI agent projects will be scrapped by 2027. The compound error table explains why.

An agentic workflow with six agent handoffs, each operating at 90% accuracy, produces 53% end-to-end reliability. Each agent call is a sequential step. Each handoff is a potential error point. The accuracy requirement does not relax because the steps are automated by agents rather than by traditional software, and the math is identical regardless of what performs the step.

This is the specific mechanism behind the projected failure rate. Agentic architectures do not solve the compound error problem, but inherit it. The architecture that works is the one that either achieves per-step accuracy above the threshold or routes low-confidence outputs to human review before they propagate, not the one that chains more autonomous steps together and hopes the errors cancel out.

"Enterprise domain does not work at 80% accuracy. Doesn't work at 90." The processes that matter, the ones with real business consequences if an error propagates, require a different class of AI. Not more prompts. Not better guardrails. A model precise enough that the math resolves in your favor. That is what execution-grade AI means in practice.

Gabriel De Dominicis, Managing Director and Head of AI at KAPTO
Gabriel De Dominicis

Gabriel is a co-founder and serves as the Managing Director and Head of AI in KAPTO. He leads KAPTO’s vision and technology. A mathematician-turned serial entrepreneur with 25+ years in enterprise IT, he focuses on execution-first AI for complex, regulated operations.

Have some questions about the topic?
Drop a message to me on LinkedIn.

Want to know more?

Explore how can KAPTO add an AI workforce to your operations.

Check out the details here

Would you like to learn more?

Unlock opportunities your business may be overlooking.
Embrace the AI revolution and operate with enterprise-level efficiency.
Together, we’ll examine real-world use cases to identify how machine learning can optimize workflows, reduce friction, and drive measurable business outcomes.