AI drive-thru accuracy: what the published numbers measure | Maple Blog

AI drive-thru accuracy: what the published numbers measure

By Maple Team · Published

What vendor, chain and study figures for AI drive-thru orders count, what the SEC found about non-intervention, and a worksheet to check your own lane.

Most published AI drive-thru figures are completion rates: the share of orders the AI finished without a person stepping in. Fewer say whether the food matched what the guest asked for. The one outside test, Intouch Insight's 2025 mystery-shop study, scored AI lanes 83% accurate against 87% across all lanes it visited.

So a vendor's 95% and a chain's 86% may not describe the same thing. Before you compare figures, find out what each one counts. Then count your own lane the same way. This guide lists the figures in print, explains what each measures, and gives you a worksheet for your own trial.

What do the published figures count?

Each row is the publisher's own claim, with its definition where it gives one. We have not tested any of these systems.

Who published itThe figureWhat it counts, per the sourceWhat to ask
Hi Auto, product page93% completion and 96% accuracy across about 1,000 storesCompletion: orders sent to the POS without staff stepping in. Accuracy: orders sent to the POS with every item and customization exactly rightDo orders its remote supervisor helped with count as completed?
Presto, home pageUp to 95% non-intervention; up to 85% automation on its Voice AI pageNot defined on either pageWhich staffing mode were those stores running?
SoundHound, drive-thru page90% AI order completion rateNot defined on the pageWhich stores and dates, and who checked the orders?
ConverseNow, platform updates95% of drive-thru orders handled end to end by AI; over 90% order accuracy (December 2024)Says the drive-thrus ran with zero remote live-agent helpHow was order accuracy checked?
Incept AI, home page95%+ AI-only completionNot defined on the pageDoes it count orders that drove off or went to the crew?
Omilia, drive-thru pageAutomates 95%+ of drive-thru ordersNot defined on the pageIs that a share of all cars, or of orders the AI started?
Wendy's, December 2023 blog86% pilot accuracy for FreshAIThe share of orders FreshAI handled without a team member stepping inWere the finished orders right?
McDonald's, September 2026 investor dayArchy takes orders in English and Spanish "with accuracy above 90%"Not defined in the remarksHow is it measured, and how often do staff step in?
Intouch Insight with QSR Magazine, 2025 study83% accuracy at AI lanes; 87% across the studyMystery shoppers placed 120 orders at AI lanes of three chains in June and July 2025 and scored themWhich vendors ran those lanes? The study does not say

Hi Auto is the only vendor here that prints separate definitions for completion and accuracy. Wendy's figure, which it calls accuracy, counts what Hi Auto calls completion.

How is completion different from accuracy?

Every car at the speaker ends in one of five ways. The AI takes the order alone and gets it right. The AI takes it alone and gets it wrong. A person helps and it ends right. A person helps and it ends wrong. Or the guest leaves without an order.

A completion rate counts the first two together. It says nothing about whether the food was right. An accuracy rate counts right orders, and it depends on who checked and which orders were in the count. The "person" can be your crew on the headset or a vendor's agent off site, and a figure that counts only your crew leaves the second group out.

The 2025 QSR Drive-Thru Report splits the AI orders this way. In 72% of them, the AI took the whole order, and those orders were 81% accurate. In 21%, the AI began and passed the order to an employee, and those were 95% accurate. The report gave three reasons for handoffs: a question the AI could not answer, a customization it did not understand, or an item out of stock. It said 62% of the wrong orders came from customization.

The same report found 34% of guests at AI lanes had to repeat their order, against 22% at other lanes. Service time was shorter at the AI lanes, 3 minutes 53 seconds against a 4 minute 15 second average in Intouch's release. Speed and accuracy moved in different directions, so track both.

What did the SEC find about "non-intervention"?

In January 2025 the SEC settled charges against Presto Automation, which sold Presto Voice. The SEC order found that Presto's reported "non-intervention" and "automated order completion" rates counted orders finished without restaurant staff stepping in. Contracted agents at off-site locations, including in the Philippines and India, still entered many of those orders. The order says a version piloted from June 2023 needed an agent to enter the order about 70% of the time through at least December 2023. Presto Automation neither admitted nor denied the findings.

Ask any vendor whether its definition leaves out people working off site. Our review of restaurant AI failures covers the Presto case with other public setbacks.

How do you measure accuracy at your own lane?

Check a sample of orders by hand, the same way in a crew-only week and in the AI trial. Record each order in a sheet like this. The first row is made up.

CarWhat the guest asked forPOS ticket matches?Food handed out matches?Who touched itWhat went wrong
1 (example)Two cheeseburger combos, one with no onion and a diet colaNo: onion left onNoAI onlyModifier
2
  1. Pick two dayparts, a rush and a steady hour, and check 25 or more orders in each.
  2. Take what the guest asked for from the speaker audio or from a crew member listening on the headset.
  3. Compare it with the POS ticket, then check a share of bags at the window.
  4. Mark who touched each order: the AI alone, your crew, or the vendor's staff. Ask the vendor for a per-order log of its own staff's help.
  5. Count guests who left without ordering as their own line.
  6. Name the error: wrong item, wrong modifier, missing item, extra item or wrong price.

How can one week produce four different numbers?

A made-up trial week shows why definitions matter. Of 200 cars that started ordering, 6 left without an order and 194 reached the POS. The AI took 150 orders alone and 138 of those were right. Vendor agents helped with 14 and 13 were right. The crew took over 30 and 28 were right.

DefinitionSumResult
Orders sent to the POS without crew help(150 + 14) ÷ 19484.5%
Orders the AI finished with no person anywhere, out of all cars that started150 ÷ 20075.0%
Right orders out of all tickets(138 + 13 + 28) ÷ 19492.3%
Right orders out of those the AI took alone138 ÷ 15092.0%

All four come from the same cars. The first counts the vendor's agents as automation. The second counts every person and every car that left. Ask each vendor which one it quotes, and report your trial on all four.

What should you ask a vendor about its numbers?

  1. What is the top and bottom of the fraction: all cars, orders started, or orders sent to the POS?
  2. Does any person off site listen to or key in orders? How often, at stores like mine?
  3. Who decides an order was right, and do they check the ticket or the food?
  4. Which stores, dates and dayparts are behind the figure?
  5. Will you report all four numbers above for my lane each week of the trial?

The one-lane pilot checklist covers the baseline and pass rules, and the drive-thru timer guide shows how to time cars by hand.

Maple publishes no drive-thru accuracy figure. Its drive-thru unit, in early access with a small number of restaurants, shows each item and the running total on its screen so guests can catch a mistake before the window, and a one-press switch returns the lane to the crew. For phone orders, the product page suggests counting correct orders, orders needing staff help and requests not completed as separate numbers.

Published by Maple, which sells phone ordering and an early-access drive-thru unit. This AI-assisted guide combines vendors' and chains' own pages, an SEC order and the 2025 Intouch Insight and QSR Magazine study with an original worksheet. The worked example is made up. It does not verify any vendor's figures or report a Maple test.