As for me, I just wondered: What percentage should fail?
Certainly not zero.
If none of your AI projects fail, you are not innovating. You are automating tasks you already understand with tools that are already proven. That may be useful, but it is not frontier work. On the other hand, if nearly everything fails, that is not innovation either. That is a bonfire with a steering committee.
While some AI projects should fail, the question is whether those failures are designed, informative, and cheap enough to be worth the learning.
Three Tiers (and Why the Number Means Nothing Without Them)
When someone delivers a 90% failure rate in a boardroom, your first response should not be panic. It should be taxonomy.
A back-office automation tool is not the same species as an initiative to reinvent the company’s operating model around autonomous agents. One is a Labrador. The other is a velociraptor. Judging them by the same standard will not produce insight.
A healthy AI portfolio contains three distinct kinds of bets, each with a different expected success rate.
The Machine is core automation: applying proven AI tools to known workflows. Summarizing tickets. Classifying documents. Flagging anomalies. The technology is mature, the process is understood, and success is expected. If your Machine projects are failing, don’t blame the algorithms. You have a data quality problem, a user adoption problem, or a governance breakdown; none of those have anything to do with the state of AI.
The Reach is adjacent innovation: extending current capabilities into new decisions or new workflows. Dynamic pricing. Demand planning. Claims triage. You understand the domain, but the project changes how decisions get made and how people work. Mixed outcomes are expected. That is not a failure of execution. That is the cost of stretching.
The Moonshot is transformational: autonomous agents redesigning major workflows, AI-native products, new business models. Uncertainty is not a bug. Venture capital lives in this territory: fund ten to hit one. If every Moonshot succeeds, the word transformational does not mean what you think it means.
These are not maturity levels; they are risk categories. A healthy portfolio contains all three simultaneously, with different funding models, different timelines, and different definitions of success. Consider customer service: a Machine project summarizes call notes after each interaction; a Reach project recommends next-best actions to agents in real time; a Moonshot redesigns the service model around autonomous agents. All three are “AI in customer service” and none should be evaluated the same way.
A single aggregate failure rate collapses all three tiers into one number. That is the innovation equivalent of saying the average person in the room is comfortable because half are freezing and half are on fire.
The Junk Drawer Problem
One reason AI failure statistics sound so alarming is that most organizations call too many things “AI projects.” A chatbot pilot is an AI project. A custom large language model trained on company documents is an AI project. A full-scale reinvention of underwriting or claims processing is also an AI project.
This is not a portfolio. This is a junk drawer. And like most junk drawers, it contains useful things, dangerous things, mystery cables, and at least one item nobody can identify but everyone is afraid to throw away.
When all of these projects are lumped together, an aggregate failure rate is meaningless and potentially actionable in the wrong direction. It gives cautious executives a reason to cut the Machines (which should be succeeding) while continuing to protect the Moonshots (which should be failing more often, more cheaply, and more informatively than they currently do).
The Failure That Actually Matters
Please note the difference between failing because the frontier is hard and failing because nobody defined success. The first is innovation. The second is theater.
An AI experiment that tests whether a model can accurately classify documents for a claims team can fail usefully. It might reveal that your historical data is an inconsistent mess, that the workflow has too many manual exceptions, or that the human review process needs a total overhaul. That is valuable learning. An AI initiative that spends nine months building a system nobody adopts, for a process nobody agreed to change, measured against benefits nobody quantified… that has not failed nobly. Congratulations on completing the corporate pilgrimage from enthusiasm to ambiguity to silence.
Does any of this sound familiar: the model works, the project fails; the tool is accurate, but employees were not trained, incentives were not aligned, or managers never made it part of how work gets done. These are not AI failures. They are management failures wearing a ChatGPT hoodie.
Understanding the failure mode defines what you should do next. A Moonshot failing on technology is a timing problem. A Machine failing on adoption is a governance problem. Neither lesson has anything to do with artificial intelligence.
Stage Gates Are Boring. Use Them Anyway.
Healthy portfolios do not let projects wander through the building until someone in Accounts Payable asks why these invoices are still arriving. Use gates. At each stage, leadership makes an explicit decision: continue, kill, hold, or redirect.
Before funding the next stage, ask what business decision or workflow the project will change. If the answer is vague, stop. Ask what would count as success and do not accept an answer that is not tied to a business outcome. Ask what would count as failure, because if nobody can define failure, the project is not an experiment. It is a hope with a purchase order. And ask what must be true for this to scale, because that is where most AI pilots die. The demo works. The economics do not. The vendor calls it enterprise-grade, which, translated from fluent Vendorian, usually means “please don’t ask any follow-up questions.”
Most importantly, every gate should record what was learned… on stone tablets mounted in the lobby. A killed project with a clear learning is not waste. A continued project with no learning is.
So, What Should Fail?
No universal law governs failure rates, but there are rules of thumb.
Machines should mostly work. If more than one in four is failing, do not blame the technology. Ask first about data quality, process ownership, and workflow integration. A high Machine failure rate is probably a foundational management problem that AI had the bad manners to make visible.
Reach projects should expect mixed outcomes. A failure rate between 40% and 60% may be entirely healthy, provided the lessons are explicit and the next project is smarter than the last. If every Reach project succeeds, you may be playing it too safe. If every one fails, you are probably skipping the Machine stage.
Moonshots should fail often, but early. A 90% failure rate is acceptable in this tier, provided those failures are cheap, fast, and informative rather than slow, expensive, and quietly swept under the boardroom rug. A failed Moonshot earned its budget if you learned something that changed what you built next.
Your Monday Morning Mandate
The next time someone presents a 90% AI failure rate, do not gasp. Ask, “Which tier?”
If 90% of your Moonshots are failing quickly, cheaply, and informatively, while the remaining 10% cure cancer and solve cold fusion, you are looking at a healthy frontier portfolio. If 90% of your Machines are failing after months of unclear ownership, weak adoption, and no workflow integration, you do not have an AI problem. You have a dysfunctional organization.
Define the three tiers explicitly. Set target success rates for each. Fund the Moonshot tier generously enough to account for expected attrition; you cannot ask for transformational bets, fund three of them, kill two, and then complain that the remaining one did not reinvent the enterprise by Thursday and claim you had a strategy.
The right failure rate is not zero. The right failure rate is the one that proves you are learning fast enough, risking intelligently enough, and not mistaking activity for progress. That may be harder to fit on a PowerPoint slide, but it is far more useful in a boardroom.
Click here for more columns from Michael Bagalman’s Data Science for Decision Makers series.
Contributor
-
View all postsMichael Bagalman is VP of Business Intelligence & Data Science at Starz and Professor of Practice at the University of Oklahoma. He has spent more than 25 years building and leading data and decision-making capabilities at organizations including AT&T, Sony, Publicis, and Deutsch. He writes the Data Science for Decision Makers column at All Things Insights and publishes Data Science Rabbit Hole on Medium. Bagalman holds degrees from Harvard and Princeton. Learn more at MichaelBagalman.com.



































































































































































































































































