Vibe Check: Should You Sign a Voice-AI Contract This Quarter?
Five voice-AI startups have raised in the last sixty days while Google and Amazon's stacks firm beneath them. A 36-month contract signed before EOY 2025 is a bet on a 2027 roadmap that hasn't been built yet. Here is the buy/wait/pilot framework I'd use this quarter.
It is the Wednesday after Labor Day and a multi-unit operator I have known for three years is sitting across from me on Sunset Boulevard with a thirty-page contract in front of her. The contract is for a voice-AI vendor I am not going to name — partly because the contract is still live, and partly because the vendor is, frankly, fine. The vendor is not the problem. The problem is the number on page seven, which is “36” — months — and the date on the cover sheet, which is the Friday after this conversation. By Friday she is supposed to sign a three-year commitment to a piece of software that did not exist eighteen months ago, in a category whose incumbents will not finish declaring themselves until at least the end of next year.
She is paying me lunch to talk her out of it, or into it, and she does not particularly care which, as long as I tell her what I would do.
So here is the contrarian thesis I want to put down on the record. Signing a 36-month voice-AI contract before the end of 2025 is a bet that the vendor’s 2027 roadmap is already frozen. It almost certainly is not. Five venture-funded startups have raised material rounds in this category in the last sixty days, two incumbent platform vendors are firming their own stacks beneath them, and the category itself is moving fast enough that what is “best of breed” in September is going to look quaint by next summer. The right move for most operators this quarter is not buy and not wait. It is pilot — on the shortest paper your procurement team can stomach.
That is the column. The rest is how I got there, and where the framework could be wrong.
What changed in the last sixty days
I have been tracking voice agents for restaurants and hotels for two years and something changed in the funding pattern in July and August. Restaurant Business, writing in late August, made the same observation in different words: the restaurant voice-AI market is, in the magazine’s framing, getting crowded. They counted a wave of startups founded since 2023 — Vox, Loman, Incept, Palona, Maple, Revmo — most of which had been heads-down through 2024 and started raising in earnest this year.
The funding side is corroborated by Crunchbase’s August venture roundup, which flagged voice AI as one of the categories pulling disproportionate capital this quarter. What both pieces note is that the technical readiness of voice agents — latency, naturalness, the ability to handle a real-world drive-thru order with modifiers — finally caught up to operator expectations sometime between mid-2024 and mid-2025. The investors I have talked to in the last month agree on that point even when they disagree on everything else.
Eric Pakravan at TenOneTen Ventures put it cleanly: “Restaurants have tried voice for years, but the AI wasn’t ready. It is now.” That is the consensus take, and it is the premise that flipped the funding switch. When a category goes from “the technology isn’t ready” to “the technology is ready,” the next twelve months always look like what we are watching happen now — whether the category is voice agents, self-driving cars, or, fifteen years ago, mobile checkout.
The most recent capital event — and the one that pushed me to sit down and write this column — was Vox AI’s $8.7 million seed round announced August 27, with the explicit pitch of building autonomous voice AI for drive-thrus and quick-service operators. An $8.7 million seed is, by 2023 standards, enormous; by the standards of this category this quarter, it is in line. The cost of capital is low here, and the next round — the Series A — is coming for several of these companies in nine to twelve months. Which means a feature war.
The buy / wait / pilot framework
This is the part of the column the operator who bought me lunch wanted me to write down. The framework has three boxes. The art is in figuring out which box you are in, and the discipline is in not lying to yourself about which box that is.
Buy — meaning sign a multi-year commitment with a single named vendor today — makes sense if and only if three things are true. One, you have a single workflow that voice can replace where the labor math is so unambiguous that a 15-20% efficiency lift would pay back the contract in under twelve months. Two, you are large enough that the vendor has dedicated implementation resources for your account and will not deprioritize you when their next-larger customer signs. And three — this is the one most operators skip — you have an honest answer to the question “what do we do if this vendor gets acquired by a platform we do not want to be on?” If you cannot answer that question without looking at your shoes, you are not in the buy box.
Wait — meaning do nothing this quarter, watch the market, revisit in six months — makes sense if you are a single-unit operator, or if your current labor model is not under acute pressure, or if your POS contract has more than eighteen months left and your POS does not have a native voice partner. Wait is also the right call if your operations team does not have the bandwidth to run a serious pilot in Q4. Voice-AI pilots fail more often from operator under-investment than from vendor under-delivery, and a half-attended pilot is worse than no pilot because it generates data you cannot trust.
Pilot — meaning sign the shortest possible paper, with the smallest possible scope, on the most easily reversed footprint — is where I think most operators above three units and below thirty actually belong this quarter. Pilot means one location, one shift band, one workflow, ninety days, with a contractual exit clause that does not require lawyers to enforce. It means you are buying information about voice agents in your specific operating environment, not buying voice agents themselves. It means your KPI is not ROI; your KPI is “did we learn enough to know what to do in March?”
Most of the operators emailing me think they are in the buy box because the vendor has convinced them they are. Most of them are actually in the pilot box. A meaningful minority — and this surprised me when I started doing the math — belong in the wait box, and the most expensive mistake those operators are making this quarter is signing because they feel like they “should be doing something about AI.”
The vendor-reported ROI problem
Every voice-AI vendor I have talked to in 2025 has a deck slide that claims a multiple — 17x ROI, 22% revenue lift, 17% labor savings — and every one of those slides has, somewhere in 8-point gray type at the bottom, a citation to a case study with one operator. I have to flag the same thing every time I write about this category because the pattern is consistent: the published ROI numbers in restaurant voice AI are vendor-reported and almost never independently audited. I do not say that to impugn the vendors. I say it because I have done enough of these case studies myself to know how the sausage gets made, and the sausage is usually one operator, one quarter, one comparison period chosen with care.
The two numbers quoted most often in operator decks this summer are Vox AI’s 17x ROI claim and Loman’s roughly 22% revenue lift / 17% labor savings claim. I have no reason to think either vendor is fabricating; I have every reason to think the underlying sample is tiny and the baseline generous. When an operator cites those numbers as evidence the category is “proven,” I push back, gently, with the same question I would ask of any vendor-reported metric: who counted the labor savings, what was the counterfactual baseline, and what time period was used? If the answer is “the vendor’s own deck,” that is data, but it is not evidence.
The good news is that the next generation of case studies — the ones that will land in Q1 and Q2 — will be run by analysts and by multi-unit operators who have done their own pilots. I would rather an operator sign a 90-day pilot now and generate their own ROI number in February than sign a 36-month deal in September based on a vendor’s number from a single operator in another state.
Hardware lock-in, POS depth, and the four risk vectors
If you are going to pilot, the question becomes which vendor — and that question is mostly a question about four risk vectors that vendors do not love to discuss on sales calls. I’ll name them in the order I weight them.
Hardware lock-in is first. Some vendors require their own purpose-built drive-thru hardware or headset stack; some run on top of what you have. The former is a steeper integration and exit cost; the latter is more flexible but often less performant in noisy environments. Neither is wrong. What is wrong is signing 36 months on hardware you have not deployed and tested in your noisiest unit. I have watched operators sign hardware-coupled deals on a quiet-room demo and discover three months in that their actual lunch-rush latency was unworkable. There is no way to learn this from a deck. Put the hardware in the worst location you have and run a Saturday.
POS integration depth is second, and it is where most early-stage voice vendors are still thin. Voice has to talk to POS, KDS, menu management, modifier hierarchies, promo logic, and — increasingly — loyalty. A voice agent that can place an order but cannot apply a loyalty discount or honor an LTO modifier hierarchy generates a stream of edge cases your shift leads triage in real time. Ask the vendor for a list of every POS field their integration writes to, and what happens when your POS pushes a menu update at 9 a.m. Vendors who cannot answer crisply have not done the work.
Multi-language coverage is third, and asymmetric across operators. If you operate in markets where a meaningful percentage of guests speak Spanish, Mandarin, Tagalog, or Korean as a first language, language coverage is not a feature — it is a requirement. Most current vendors will say they support multi-language; in practice that means English is production-grade and the second language is “in pilot.” Ask for evaluation scores by language with a baseline date, and which language is next on the roadmap. If the vendor cannot give you a date, the answer is functionally “we do not have a roadmap.”
Latency and “human-in-the-loop” disclosure is fourth, and the one I would press hardest on. The category went through a credibility crisis last year when Presto disclosed in SEC filings that a substantial percentage of its voice interactions were actually being handled by remote human agents rather than autonomous AI. The disclosure itself wasn’t the scandal — every vendor uses some human escalation for low-confidence interactions, and that is appropriate operational design. The scandal was that operators did not know. If your vendor will not disclose, contractually and in writing, the percentage of interactions that are AI-only versus human-assisted, and will not commit to a monthly transparency report on that ratio, you have a vendor problem before you have a technology problem. Ask on the first call. The answer tells you most of what you need to know.
What about Google and Amazon
The reason I keep telling operators to pilot rather than buy is not that I think the startups are bad. Several are, in my opinion, very good. The reason is that incumbent platform vendors — Google, Amazon, increasingly Microsoft — are firming their own stacks underneath this category, and the contours of what an “AI voice agent for restaurants” looks like in 2027 will be shaped at least as much by what Google does with its real-time speech APIs as by what any of the named startups ships. The directional claim: a category dominated by startups in 2025 is one where the infrastructure is owned by hyperscalers by 2027, and the startups that survive will be the ones whose product is the workflow layer — menu logic, POS integration, operator dashboards — rather than the speech-recognition layer itself.
If that is right, signing a 36-month contract today on a startup whose core differentiation is its speech model is signing on the wrong moat. The right moat — workflow depth, integration breadth, operator dashboards — is exactly what the startups are still building. In twelve months you will be able to see, with reasonable clarity, which of these vendors built the right thing. In September 2025 you are being asked to guess.
What 36 months actually buys you
Here is the math I would put in front of any operator considering a multi-year voice deal. A 36-month contract typically saves you 15-25% off rack rate on per-location monthly fees. On a stack pricing $400-$700 per location per month for QSR, that is meaningful across a multi-unit footprint. Call it $30-$80 per location per month in committed-discount savings.
What you are paying for that savings is the right to exit — not in the sense of an early-termination fee, which most operators model correctly, but in the opportunity cost of being locked into a vendor whose 2027 roadmap may diverge from where the category goes. If voice-agent capability is doubling every twelve months — and the funding flow suggests it is — the vendor you signed in September 2025 is, by September 2027, two generations of capability behind whatever emerged in the meantime. The committed-discount math saves $30-$80 per location per month. The capability-gap math, if the category moves the way I think it will, costs somewhere between 5% and 15% of the revenue lift you could have captured on a more current platform.
I have not seen an operator do that math out loud in a procurement meeting. I have seen many operators do the committed-discount math out loud. The asymmetry of which math gets done in the room is most of why 36-month deals get signed too often in fast-moving categories. The discount is visible on a spreadsheet. The opportunity cost is not.
How I would structure a Q4 pilot
If you are in the pilot box, here is the structure I would use. Six-month contract, not three years. One location, picked deliberately to be your median location rather than your best one, because best-location data does not generalize. One workflow only, ideally the highest-volume one (drive-thru lunch rush for QSR, inbound reservations and call-ahead for casual dining). Defined success metrics signed by both sides before kickoff, with a 90-day midpoint review and a contractual right to terminate if midpoint metrics are not hit.
The vendor will push back on the six-month length. The pushback is the test. Vendors confident in their product will agree to a six-month pilot because the renewal economics are good. Vendors who insist on 24 or 36 months are telling you something about either their internal forecast or their anxiety about the competitive landscape. Listen. The pushback is data.
On pricing, do not accept “pilot pricing equals rack rate.” Both parties are sharing the risk of a category still proving itself. Pilot pricing should be 60-75% of rack rate, with a contractually defined uplift on renewal. Price the pilot as if you are selling the vendor a market-research asset, because you are.
Where I could be wrong
Three places this thesis could be wrong, and they are worth naming because the operator across the lunch counter asked me to name them.
One: the category could consolidate faster than I think. If one of the named startups raises a large Series A in Q4 and uses it to roll up two or three competitors, the vendor concentration risk I am flagging gets more acute, not less, and the case for picking the winner now gets stronger. I do not currently see signals of that consolidation happening this year, but I would not bet against it for 2026. The forthcoming framework piece on voice agent maturity goes deeper on the consolidation scenarios than I can here.
Two: incumbent POS vendors could pre-empt the standalone voice market by shipping good-enough native voice in 2026, which would compress the pricing the standalone vendors can charge and make a 36-month deal signed this quarter look genuinely overpriced in retrospect. I think this is the most likely scenario, actually, and it is one of the strongest arguments against signing now.
Three: there is a class of operator — large enterprise QSR, where the labor math is unambiguous and the integration cost is going to be enormous regardless of when you sign — for whom the buy box is genuinely the right answer this quarter, and the wait/pilot framework is wrong. The upcoming piece on McDonald’s AI drive-thru goes into how the very largest operators are thinking about this differently, and the short version is that at sufficient scale the calculus changes. If you are running fewer than fifty units, that is not you.
What I told her
The operator who bought me lunch did not sign the contract on Friday. She wrote the vendor a polite email proposing a six-month pilot at 70% of rack rate on her median-volume location, with a 90-day midpoint review and contractual termination rights if the AI-only interaction percentage drops below a number we wrote into the email. The vendor came back on Tuesday with a counter-offer that was, to my surprise, more generous than I expected. They wanted the operator badly enough to give her most of what she asked for. Which tells you something about which side of this market actually has leverage right now.
I do not know yet whether the pilot will work. I do know that the cost of being wrong on a six-month pilot is small, and the cost of being wrong on a three-year contract is large, and in a category moving as fast as this one, the asymmetry of those two costs is the entire game. Buy if the math is unambiguous and you can answer the acquisition question without looking at your shoes. Wait if you are not ready. Pilot if you are anywhere in between. Sign short paper. Make the vendor share the risk. Revisit in March.
That is the framework. The contract on the counter at lunch had thirty pages. The framework has three boxes. Most of what we are paying these vendors for, in 2025, is the information about which box we are actually in. The discipline is being honest about it.
— Sofia leads Vibe Check vendor reviews for TableTransfers. Tips: vendors@tabletransfers.com.
The Voice Agent Maturity Curve
mise
·12 min read
The Four Margins of a Restaurant
mise
·14 min read
The AI Premium in Hospitality M&A: Broker Story or Real Number?
the bottom line
·9 min read
What the DoorDash/SevenRooms Deal Actually Buys
the bottom line
·11 min read
Related posts
vibe check
·16 min read
Desk Review: Lightspeed Restaurant, the Quiet Half of the Duopoly
vibe check
·18 min read
Desk Review: OpenTable's 'System of Record' — what restaurants are actually agreeing to on April 16
vibe check
·18 min read