Skip to content
All Insights
§ 06 — Insights

Anatomy of a Failed Enterprise AI Rollout: Five Recurring Failure Modes

Klarna, McDonald's, Air Canada, Zillow, and IBM Watson Health each produced high-visibility AI reversals. The individual stories differ. The evidentiary gap they share is the same.

July 10, 20268 min read
§ SHARE

A short catalogue of publicly reported enterprise AI reversals is, by now, familiar to any restructuring or audit committee. The names change — Klarna reversing customer-service automation, McDonald's discontinuing a multi-year drive-thru voice pilot with IBM, Air Canada being held to a chatbot's inaccurate refund policy, Zillow winding down its algorithmic home-buying arm, IBM Watson Health being sold off in pieces — but the pattern beneath the individual stories is consistent.

The pattern is not that AI failed to perform. It is that the enterprise could not produce, at the point of reversal, a record capable of explaining what had actually been in production, what it had produced, at what cost, and who owned the decision. What follows is five recurring failure modes, drawn from the public reporting.

1. The vendor promise outran the evidentiary record

In the McDonald's–IBM voice ordering pilot, three years of investment across roughly one hundred stores ended in a discontinuation announcement. What is not visible from the outside is whether the organization retained the operating data required to distinguish model performance from store-level integration issues, or to defend the discontinuation against the vendor's counter-narrative. Where that record does not exist, the reversal cannot be cleanly negotiated.

2. The system in production was not the system that was approved

Klarna's public reversal of its AI-first customer-service posture illustrates a different variant. Between initial deployment and the reversal, the underlying assistant went through multiple model swaps, prompt revisions, escalation logic changes, and integration expansions. Reconstructing which version produced which customer outcome, on which date, is a nontrivial exercise even for a company that was actively publicizing its AI results.

3. The legal exposure was inherited without the contract to support it

Air Canada's tribunal loss over a chatbot bereavement-fare policy that the bot had fabricated is often read as a warning about hallucinations. It is more precisely a warning about the ownership of an AI system's outputs. The airline argued, unsuccessfully, that the chatbot was a separate entity. The tribunal treated the outputs as the company's own. Any AI system deployed to customers now carries the same treatment risk, and the contract with the vendor is either the shield or the exposure.

4. The economics were never modeled against real-world data drift

Zillow's iBuyer arm, Zillow Offers, was wound down in 2021 after the algorithmic pricing model underperformed against actual market conditions. Reported losses were in the hundreds of millions. The failure was not that the model was built badly. The failure was that the operating assumption — that an algorithm calibrated on historic transactions would remain calibrated in a rapidly moving market — was never audited against the incoming data as the market moved.

5. The strategic asset became unsellable because it could not be described

IBM's divestiture of Watson Health assets closed a chapter that had absorbed billions in acquisitions, integration, and marketing. What the sale process revealed is that a fragmented AI asset — spread across product lines, contracts, data-use agreements, and vendor dependencies — is materially harder to price and transfer than an equivalent conventional software asset. Buyers discount aggressively for anything they cannot reconstruct.

The shared underlying condition

In each of these cases, the operational failure is only the surface. The condition that made the failure unrecoverable — or recoverable only at a heavy discount — is the absence of a reconstructable record. The evidence recoverability determines the workout options. The workout options determine what value, if any, can be preserved. The provenance problem is the same pattern seen from the vendor side; the disposition framework is the response to it.

  • Contracts that no one has re-read since signature.
  • Model or prompt changes made by the vendor without notice.
  • Outputs shipped to customers without a policy owner.
  • Model calibration assumptions that were never re-tested.
  • Portfolio assets whose scope cannot be described to a buyer.
§ SHARE
RECOVERY IMPLICATION

Publicly reported AI failures cluster around evidentiary gaps, not model quality. Where the record cannot support the reversal, the workout options collapse to write-off.

Sources & References
§ RELATED

Continue reading

Next in sequence
The $30–40 Billion ROI Gap: Why Enterprise AI Spend Isn't Producing Returns