The Voice-AI Scorecard: What SoundHound's 14,000-Location Footprint Tells Us

Sofia Reyes building a voice-AI deployment scorecard from SoundHound's Q3 disclosures

SoundHound named seven major chains on its Q3 call and posted 68% revenue growth. I used those disclosures to build the first operator-facing scorecard ranking voice-AI deployment maturity against the hype curve.

It is 7:14 a.m. on a Wednesday and I have three browser windows open: SoundHound’s Q3 GlobeNewswire release, the SEC 8-K earnings exhibit, and the Motley Fool transcript of the November 7 call. The fourth window is a spreadsheet I have been threatening to build for six months: a vendor-agnostic scorecard for voice-AI deployment maturity in restaurants. The trigger to finally build it was not a vendor briefing. It was SoundHound’s willingness, for the first time, to name names — MOD Pizza, Habit Burger, Red Lobster, Torchy’s, Chipotle, Casey’s, Peet’s — at a level of specificity the category has been avoiding for two years.

Here is the contrarian thesis I want to defend in this column: the voice-AI category does not have a technology problem, it has a disclosure problem, and SoundHound’s Q3 named-chain list is the first dataset rich enough to separate deployment maturity from deck-deep marketing. The companies that look strongest in the trade press are not the companies running the most lanes. The companies running the most lanes are mostly invisible because their operators will not let them publish case studies. SoundHound’s Q3 call cracked that pattern open, and once you have a single vendor talking publicly about a 14,000-location footprint with seven named anchor brands, you can finally back-solve a scorecard the rest of the category has to be measured against.

I will not pretend the scorecard I am about to publish is perfect. It is a first cut, with weights I expect to revise after I run it past two operators and one investor. But it is sharper than what currently passes for analysis in this category, which is mostly press-release stenography and analyst price targets that move on tweets. If you are an operator trying to decide whether to pilot voice ordering in Q1, this is the framework I would use, and SoundHound is the natural anchor case because they are the only one currently disclosing at this granularity.

What SoundHound actually said, and why the granularity matters

Let me be clear about what is new in the Q3 disclosure, because the headline number — $42M revenue, up 68% year-over-year, full-year guide raised to $165–180M per the GlobeNewswire release — is not really the news. Growth at that rate is impressive but it is also the kind of number every voice-AI vendor will claim by Q2 next year because the comparison base is small enough that almost any contract win moves the needle. What matters is the named-chain list and the franchise-system disclosures.

On the November 7 call, per the Motley Fool transcript, management walked through seven anchor accounts — MOD Pizza, Habit Burger, Red Lobster, Torchy’s Tacos, Chipotle, Casey’s, and Peet’s Coffee — and added three franchise-system wins where individual franchisees are deploying ahead of corporate mandates: Firehouse Subs, Five Guys, and McAllister’s Deli. Some of these are voice ordering in the drive-thru. Some are phone-channel voice agents. Some are agentic ordering that wraps voice plus chat. The specificity is what changes the analytical job. Before this call, you had to triangulate between press leaks, LinkedIn job postings, and the occasional franchisee podcast appearance to figure out who was actually deploying at scale. Now you have a vendor willing to anchor the conversation in named-brand reality.

The disclosure is also useful because it tells you what SoundHound is not claiming. They are not claiming Chick-fil-A. They are not claiming Wendy’s beyond what was already public. They are not claiming Domino’s. The absence of those names is informative — it tells you which contests are still genuinely open and which competitors (Yum’s internal stack, Google’s restaurant builds, OpenAI’s enterprise pilots) are credible enough that SoundHound has not displaced them at the brand level.

That is exactly the kind of negative-space disclosure that lets you build a scorecard. You need to know what a vendor will and will not say in writing under SEC scrutiny. Anything they would say in a deck but not on an earnings call you should weight at roughly zero.

The scorecard, version 0.1

Here are the seven dimensions I am scoring vendors on. I will run SoundHound through them as the anchor case because they are the only vendor with enough public disclosure to evaluate honestly. For the other vendors in this category, I will leave the scores as ranges until they publish at comparable granularity.

1. Named-brand depth (weight: 20%) — How many enterprise brands has the vendor put in writing in a regulated filing or earnings call? Decks do not count. LinkedIn screenshots do not count. The bar is a name in a 10-Q, 8-K, or earnings transcript.

2. Channel breadth (weight: 15%) — Does the deployment cover drive-thru voice, phone-channel voice, in-store kiosk voice, and agentic ordering, or just one of those? Single-channel vendors are not wrong, but they are not yet platform plays.

3. Franchise-system penetration (weight: 15%) — Are individual franchisees deploying ahead of corporate, or does the vendor only sell to corporate? Franchisee-led adoption is a much stronger durability signal because franchisees are paying out of unit economics, not test budgets.

4. Operator-side transparency (weight: 15%) — Does the operator (not just the vendor) talk publicly about the deployment? If the operator will not say the vendor’s name on a conference panel, treat the case study as soft.

5. Failure-mode disclosure (weight: 10%) — Has the vendor or operator described, in writing, what went wrong on the way to deployment? Vendors that only publish wins are operating below the maturity threshold I care about.

6. Unit-economics defensibility (weight: 15%) — Has the deployment been described in terms an operator can underwrite — orders per labor hour saved, accuracy lift, AOV lift, or comparable metrics — or is the entire story aspirational?

7. Multi-language and regional fit (weight: 10%) — Does the system handle Spanish in U.S. drive-thrus, French in Quebec, and the dialect breadth of an actual hospitality workforce, or is it English-only with a Spanish coming-soon page?

I will say upfront that the weights are mine, not gospel. An operator who is single-channel and English-only and franchisee-driven will weight these differently. The framework is the contribution; the weights are the conversation.

SoundHound against the scorecard

Running SoundHound through this framework gives a composite that is honestly higher than I expected when I started the exercise. Let me walk through it dimension by dimension because the dimensional scores matter more than the composite.

Named-brand depth: 9 of 10. Seven anchor brands plus three franchise-system disclosures, in writing, on a regulated earnings call. This is the highest score in the category right now, and it is not particularly close. The only way this drops is if a competitor lands a Chick-fil-A or Wendy’s enterprise mandate and discloses it at the same granularity, which is possible but not yet on the public record. I want to flag, though, that “named in a transcript” is not the same as “deployed at every location.” A brand can appear on this list because of a pilot in a single market. We do not yet have per-brand location counts, and we should ask for them.

Channel breadth: 8 of 10. The Q3 disclosures cover drive-thru voice (Casey’s, Habit, MOD), phone-channel agents (Red Lobster, Peet’s), and agentic ordering wrappers (Torchy’s, Chipotle to varying degrees). That is genuine multi-channel coverage, which is rare in the category. The reason this is not a 10 is that kiosk voice is still thin in the public disclosures, and kiosk is where some of the most defensible unit economics live for QSR.

Franchise-system penetration: 9 of 10. Three franchise-system wins (Firehouse, Five Guys, McAllister’s) where franchisees are deploying ahead of corporate is the cleanest durability signal you can ask for. Franchisees do not pay six figures for vendor pilots out of marketing budget; they pay out of unit P&L, which means the deployment is being justified on labor or throughput math that survives an operator’s underwriting standards. I am holding back one point because we do not yet know whether franchisees are renewing, and renewal is the actual durability test.

Operator-side transparency: 6 of 10. This is where SoundHound’s disclosure outruns the operators. The brands have not all been equally vocal. Casey’s has been the most willing to talk in operator forums, and MOD has done some panel work, but Chipotle and Red Lobster have been quiet on the operator side. That is fine for a Q3 disclosure but it caps the scorecard until the operator community is willing to corroborate at panel events. I would expect this to improve in the spring conference cycle.

Failure-mode disclosure: 5 of 10. SoundHound has been more willing than most to talk about edge cases, but the public record on what has not worked is still thin. The category as a whole scores poorly here; SoundHound is at the higher end of a low bar. The forthcoming May framework piece in our voice-agent maturity series will spend more time on failure-mode taxonomy, because right now we are all relying on operator backchannel for what went wrong, and that is not a scalable analytical method.

Unit-economics defensibility: 7 of 10. The Q3 call gestured at labor and throughput economics, but the public metrics are still framed at the vendor level (subscription revenue, royalty revenue) rather than the operator level (orders per hour, accuracy, AOV lift, labor offset). The franchise-system wins are the strongest implicit unit-economics signal because franchisees do not buy on faith. I would like to see an explicit per-brand AOV or accuracy disclosure before this score moves up.

Multi-language and regional fit: 8 of 10. SoundHound’s heritage in conversational voice gives them genuine multilingual capability, and the Casey’s and Torchy’s deployments cover regions where Spanish-language drive-thru is not optional. This is one of the dimensions where the company’s pre-restaurant history actually matters; voice-first incumbents have a real advantage over teams that started from text and bolted on speech.

The composite, applying the weights above, lands at 7.5 out of 10. That is a strong score, and I want to be honest that it is higher than I expected when I started. The reason is that the named-brand depth and franchise-system penetration dimensions are weighted heavily, and SoundHound is genuinely best-in-category on both right now. If I shift the weights toward operator-side transparency and failure-mode disclosure — which I might, after some operator conversations — the composite drops into the 6.5 to 7 range, which is still the highest in the category but less of an outlier.

The honest read is that SoundHound is the only voice-AI vendor currently scoring above 6 on this framework, and the gap is largely a disclosure gap rather than a capability gap. There are competitors with technically comparable systems whose operators will not let them publish, and the operators will not let them publish because the deployments are too early or too uneven. That is the hype-curve signal: the category is further along in named-brand acquisition than the public record suggests, and SoundHound’s willingness to disclose is dragging the rest of the field toward a transparency standard they would prefer to avoid.

What the disclosure pattern tells us about competitive position

The thing the Q3 disclosure does not say is sometimes louder than what it does say. SoundHound did not claim McDonald’s, Yum brands, Chick-fil-A, Wendy’s beyond what was already public, Burger King in the U.S., or Domino’s. Those omissions tell you the contests still in play. The upcoming May piece on McDonald’s drive-thru AI will dig deeper into the McDonald’s-specific situation, but the short version is that McDonald’s is the prize that has been most contested across vendors, and SoundHound’s silence on it is consistent with the public reporting that Google’s restaurant team has been the primary partner there.

The contested-brand pattern is interesting for a second reason. The brands SoundHound did name skew toward the mid-cap and regional-strength end of the spectrum — Casey’s is dominant in the Midwest but not coastal, MOD and Habit are strong regional concepts, Torchy’s is a Texas-anchored concept with national ambitions. The Chipotle and Red Lobster names are the largest by location count, but Chipotle’s deployment is at the agentic-ordering layer rather than the drive-thru voice layer, and Red Lobster’s is a phone-channel agent for reservation and order capture. The pattern looks like SoundHound winning by going deep in mid-market and franchise-system accounts while the tier-one mega-brands remain contested.

For an operator reading this scorecard, that pattern matters because it tells you where SoundHound is most likely to have the best-trained acoustic models, the most operator-tested integrations, and the deepest customer-success bench. If you are a 200-unit regional concept evaluating voice-AI vendors in Q1, SoundHound’s mid-market depth is exactly the kind of evidence you want. If you are a 5,000-unit national chain, you should weight more heavily the franchise-system signals, because that is where SoundHound has shown they can scale operationally.

I should also flag the SoundHound Pass announcement from November 6 — the agentic ordering layer that wraps multiple channels — because the Q3 disclosure has to be read alongside that product expansion. Pass is the vendor-side bet that the voice-AI platform layer is going to be defined by orchestration across channels rather than by raw speech recognition accuracy. That bet is consistent with the channel-breadth disclosure pattern and consistent with the franchise-system economics. It is the right bet to be making in late 2025, and the scorecard above is partially designed to test whether competitors are making the same bet or are still pricing single-channel deployments.

Where the scorecard breaks, and where I want operator feedback

I want to be honest about where this framework is weakest. The largest gap is that I am not yet able to score per-deployment quality. Two vendors can each have eight named brands and ten franchise systems and have completely different deployment quality, and the scorecard as written does not distinguish them. The way to fix that is operator-side telemetry — order accuracy, escalation rates, customer-experience scores — but those metrics are not yet publicly disclosed by any vendor or operator at meaningful granularity. Until they are, the scorecard will reward vendors who disclose breadth over vendors who run deeper but quieter deployments. I am aware of that limitation and I think it is the right tradeoff given that the alternative — relying on vendor decks — is worse.

The second gap is that I am not yet scoring the data-ownership question. Some voice-AI deployments leave the operator with usable conversation data they can repurpose for marketing, training, or BI. Others lock the data into vendor-side systems. That distinction is going to matter enormously over the next two years, but the public disclosures do not yet support a scoring dimension on it. I expect to add a data-rights dimension in the next version, probably at a 10% weight that comes out of the failure-mode and operator-transparency buckets.

The third gap is that the scorecard does not yet integrate the labor-side picture. Voice AI in restaurants is partly a labor-replacement story and partly a labor-augmentation story, and the framework I have built does not differentiate. A vendor whose deployments shift labor from order-taking to hospitality is doing a very different job than a vendor whose deployments eliminate labor at the order point. Operators care about both, but they care about them in different ways depending on their format and brand positioning. I would like to add a labor-model dimension, but I want operator input on whether to weight it as a hard score or as a qualifier on top of the composite.

If you are an operator running voice AI in production, or piloting it seriously in Q1, I want to hear from you. The scorecard will get better with operator feedback faster than it will get better with vendor feedback. Vendors will tell me my weights are wrong in self-serving ways. Operators will tell me what the framework is missing in ways that compound.

What I would do with this scorecard if I ran a 500-unit chain

The point of the scorecard is to make pilot decisions easier, so let me end with the operator playbook I would run if I were a VP of operations at a 500-unit regional chain looking at voice-AI in Q1.

First, I would use the framework to filter the vendor short list down to three. Anything scoring below 5 on named-brand depth comes off the list immediately, because it means the vendor has not yet been willing to defend a public-record deployment claim. That is a hard filter and it is more aggressive than what the category currently practices. I am comfortable being more aggressive because the alternative is letting vendors hide behind decks.

Second, I would weight franchise-system penetration higher than corporate-led deployments in my own evaluation, regardless of whether I am a franchise or corporate-owned organization. Franchisee-led adoption tells you the deployment survives operator-level underwriting. Corporate-led adoption can be a real signal too, but it can also be a brand-marketing decision, and you cannot tell which from the outside. Franchisee-led is harder to fake.

Third, I would demand operator references at the GM level, not the corporate-marketing level. The vendor’s case-study deck will be sourced from a corporate stakeholder who selected them. The GM running the lane every day will tell you what actually happens at 11:50 a.m. on a Saturday when the system is under load and the headset cuts out. If a vendor will not put you on the phone with three GMs running their system, treat that as a 3-point deduction on operator-side transparency.

Fourth, I would ask the vendor to commit, in the master service agreement, to a failure-mode disclosure obligation. The vendor has to tell you, in writing, when a peer deployment has experienced a material performance issue. That contract clause is unusual but not unreasonable, and it shifts the failure-mode disclosure score from a hope to an obligation. Vendors who will not sign that clause are telling you something important about the category and about their internal confidence.

Fifth, and most importantly, I would pilot two vendors in parallel for at least one quarter before committing to a system-wide rollout. The category is moving fast enough that single-vendor pilots are a bad bet right now. Run two side-by-side, in comparable markets, with shared evaluation criteria, and you will learn more in 90 days than the category has published in 24 months. The pilot cost is real but it is small relative to the cost of a wrong-vendor rollout across 500 units.

The scorecard I have built above is version 0.1 and I expect to publish v0.2 in the spring after operator feedback. For now, the headline read is: SoundHound is the only vendor currently scoring above 6 on this framework, the gap is more disclosure than capability, and the right way to read the Q3 earnings call is as a signal that the disclosure floor in the voice-AI category just moved. Competitors will have to match SoundHound’s named-brand granularity or accept that their composite scores will lag, not because their systems are worse, but because they are still hiding behind decks. The operators reading this column are the ones who get to decide whether that disclosure standard sticks. I think it should. I think it will. I think the next four quarters are going to make this scorecard a lot more crowded at the top.

— Sofia leads Vibe Check vendor reviews for TableTransfers. Tips: vendors@tabletransfers.com.

Featured More

The Voice Agent Maturity Curve

mise

·

12 min read

The Four Margins of a Restaurant

mise

·

14 min read

The AI Premium in Hospitality M&A: Broker Story or Real Number?

the bottom line

·

9 min read

What the DoorDash/SevenRooms Deal Actually Buys

the bottom line

·

11 min read

Browse all 494 posts

Related posts

Desk Review: Lightspeed Restaurant, the Quiet Half of the Duopoly

vibe check

·

16 min read

Desk Review: Lightspeed Restaurant, the Quiet Half of the Duopoly

Desk Review: OpenTable's 'System of Record' — what restaurants are actually agreeing to on April 16

vibe check

·

18 min read

Desk Review: OpenTable's 'System of Record' — what restaurants are actually agreeing to on April 16

Desk Review: Toast Drive-Thru — the bundle, the moat, and the 15-unit floor

vibe check

·

18 min read

Desk Review: Toast Drive-Thru — the bundle, the moat, and the 15-unit floor