← The Atlas Blog

McDonald's IBM Drive-Thru

Pilot 2021–2024 · Termination announced June 2024

Voice AI pilot terminated after two years across 100+ locations due to systematic order failures.

Purpose Reliability Intelligence

What Happened

McDonald's was supposed to revolutionize the drive-thru.

IBM's Automated Order Taking system — an AI voice model deployed across more than 100 McDonald's locations — was designed to take customer orders without human involvement. The promise: faster service, reduced labor costs, consistent order accuracy.

The reality, documented across two years of customer videos, social media posts, and franchise operator feedback, was different.

The AI added items customers didn't order. It failed to understand accents. It confused similar-sounding menu items. It generated orders that bore little resemblance to what customers had said. Viral videos of the system confidently ordering dozens of chicken nuggets, adding unwanted ice cream, or simply failing to process basic requests became a recurring feature of AI failure coverage in 2023 and 2024.

In June 2024, Chief Restaurant Officer Mason Smoot sent an internal memorandum to franchise operators announcing the termination of the IBM partnership. McDonald's would end the AI drive-thru pilot.

The Atlas Analysis

Purpose ≈ 45/100 — Level 2

The drive-thru AI existed to reduce labor costs and improve throughput. These are legitimate operational goals. The Purpose failure is that the system was deployed at scale — 100+ locations serving real customers — without evidence that it could reliably achieve its primary function: accurately taking orders.

Signal #3 — "Will it solve the problem in the company?" The system did not reliably take accurate orders. The problem it was supposed to solve was not solved.
Signal #10 — "What would a successful system look like?" A system that adds items the customer didn't order, fails on accents, and generates viral failure videos is not a successful drive-thru AI by any measurable standard.

The deeper Purpose question is whether a drive-thru AI should have been deployed across 100 locations before the core function was proven reliable. Pilots exist to validate purpose before scale. In this case, scale preceded validation.

Reliability ≈ 30/100 — Level 1

Reliability asks: does the system work consistently? The answer for the McDonald's IBM system is documented and unambiguous: no. Across 100 locations and two years, the system produced enough failures to generate a sustained pattern of viral documentation and franchise operator complaints sufficient to trigger termination of a partnership with one of the world's largest technology companies.

Signal #45 — "How often does the AI produce incorrect outputs?" The failure rate was high enough to be systematically visible through customer-generated content and to motivate a franchise-wide termination decision.
Signal #44 — "Does the AI perform consistently across different users?" The accent failures specifically indicate that the system performed materially differently for different customer demographics. A system that works for some users and fails for others is not reliable — it is selectively functional.
Intelligence ≈ 38/100 — Level 2

The system's failures were primarily at the level of basic speech recognition and order comprehension — the core intelligence task it was designed to perform. Adding unrequested items, confusing similar menu items, and failing to process standard orders represents Intelligence failure at the function level.

Signal #29 — "Will someone verify the results at the end?" In a drive-thru context, the "verification" is at the pickup window — after the order has been placed, charged, and prepared. By the time the error is visible to the human, the kitchen has already processed an incorrect order. The verification loop is too slow to prevent the failure.

Signals That Would Have Caught It

01

No demographic performance validation before deployment. Accent and dialect variation in voice AI is a documented and predictable challenge. A pre-deployment evaluation across the demographic diversity of McDonald's customer base would have identified this failure mode before a single real customer ordered through it.

02

No operational error rate threshold for deployment approval. The system was deployed across 100 locations without a defined acceptable error rate. Every order that adds unrequested items, generates incorrect totals, or fails to process represents a measurable failure. That rate should have been defined and validated before scale.

03

No customer feedback loop integrated into performance monitoring. The failures were visible through customer videos and social media before they were visible in internal reporting. A system operating at 100 locations should have real-time visibility into order accuracy rates — not rely on viral documentation to surface a systematic problem.

What It Cost

Two years of deployment across 100+ locations before termination. The franchise operator cost — in incorrect orders, customer complaints, and remediation — was borne by individual operators throughout the pilot period. The reputational cost was borne by both McDonald's and IBM.

The partnership termination was itself a cost signal: IBM had staked significant enterprise AI credibility on this deployment. The announcement that McDonald's was terminating the pilot sent a clear market signal about the gap between voice AI capability claims and operational reality in high-volume, high-diversity customer environments.

The Lesson

A system that works in a controlled evaluation environment and fails in production is not evidence that the use case is wrong. It is evidence that the deployment decision was premature.

This case illustrates the Purpose pillar's most important secondary question: not just "should this AI exist?" but "is this AI ready to exist at this scale?" Scale should follow validated reliability, not precede it.

The McDonald's case is also a Reliability lesson about demographic coverage. A voice AI that fails on accents is not failing randomly — it is failing systematically for specific populations. Systematic demographic failure is not a reliability edge case. It is a bias finding with a Reliability consequence.

When AI deployment decisions are made at the scale of 100+ locations, the customers who experience the failures are not test users. They are real people placing real orders. The cost of Reliability failure at that scale is borne by customers before it is visible to the organization deploying the system.

References

  1. Internal franchise memorandum, Chief Restaurant Officer Mason Smoot, June 2024.
  2. CNBC / Restaurant Dive: "McDonald's ends IBM drive-thru voice order test," June 2024.
  3. Associated Press / The Hindu: "McDonald's is ending its test run of AI-powered drive-thrus with IBM," June 2024.

FREE · 15 MINUTES

Book a free Atlas Readiness Review

Book Your Review →