Carbuki Insights
AI Agents Finish 32.5% of Complex Real-World Tasks. Here Is What That Means for Your Phones
Source: Cox Automotive, 2025 Fixed Ops and Ownership Study and 2025 Car Buyer Journey Study, published April 2026.
Every AI phone demo for dealers goes the same way. A calm synthetic voice books a service appointment, fields a trade question, absorbs one objection, and transfers cleanly to a manager. Everyone nods. Nobody asks the follow-up that actually matters: what share of your real calls look like that one?
A benchmark accepted to ICLR 2026 offers the closest thing the industry has to an honest answer. It is not the number the demo implies, and it is more useful than the number the demo implies.
The 32.5% number
VitaBench, built by Meituan's LongCat team, drops AI agents into what its authors describe as the most complex life-services simulation published to date: 66 tools spanning food delivery, in-store commerce, and travel booking, with tasks derived from real user requests. To pass, an agent has to reason across time and place, chain unfamiliar tool calls, proactively clarify ambiguous instructions, and track a customer whose intent shifts partway through a multi-turn conversation.
That is a fair description of a Tuesday morning in a dealership BDC.
The results are sobering, and specific. Across 100 cross-scenario tasks - the ones requiring an agent to operate in more than one domain at once - even the most advanced models completed only 32.5% successfully. On the 300 simpler single-scenario tasks, success rates stayed below 62% (VitaBench, ICLR 2026).
Myth vs. data
- The pitch: a modern AI agent can handle any customer conversation end to end.
- The benchmark: on tasks built from real consumer service requests, frontier models finished 32.5% of multi-domain tasks and under 62% of single-domain ones.
- The read: capability is real but uneven, and the variable that best predicts success is not which model you bought. It is how tightly the task was scoped before you handed it over.
Complexity, not capability, is the binding constraint
Notice the shape of the finding. The same models, evaluated the same way, nearly doubled their success rate when the task stayed inside one domain instead of crossing several. That is not a story about AI being unready. It is a story about task design.
The benchmark's authors name the failure mode directly: these tasks require agents to proactively clarify ambiguous instructions, and that is precisely where agents underperform. When critical information is missing, an agent under pressure to be helpful tends to act rather than ask. In a chat window, that produces a wrong reservation. On a dealership phone line, it produces a payment quote nobody in the store agreed to.
For dealers, this reframes the buying question. Stop asking whether the AI is good enough yet. Start asking which of your calls are single-scenario and which are cross-scenario.
The calls that are cross-scenario
Some inbound calls are genuinely hard, and the market has made them harder. Edmunds reported that in Q2 2026, 29.6% of trade-ins toward new-vehicle purchases carried negative equity, with the average underwater amount reaching $6,884, the highest second-quarter figure on record. Buyers rolling that debt forward paid an average of $944 per month, versus a $777 industry average, and are projected to pay $16,270 in lifetime interest against $9,811 for the average new-vehicle buyer (Edmunds, July 2026).
A caller in that position is asking one question out loud and four underneath it: what is my car worth, what do I still owe, what can I qualify for, and what does that do to my payment. Valuation, payoff, credit, and desking are four domains. That call is the benchmark's 32.5% case - and it is also the call where a confident wrong answer costs you the deal and possibly a compliance headache.
| Call type | Domains involved | Realistic scope for an AI agent |
|---|---|---|
| Service appointment for a known customer | Scheduling | Full handling, including booking |
| Hours, directions, part availability, RO status | One system lookup | Full handling |
| First-service booking after delivery | Scheduling plus CRM record | Full handling with confirmation |
| Sales inquiry on a specific stock number | Inventory plus availability | Qualify, capture, book the appointment |
| Trade valuation with an open loan | Valuation, payoff, credit, desking | Gather context, warm transfer |
| Payment or approval questions | Credit, lender programs, F&I | Gather context, warm transfer |
The right design is not AI versus human. It is an agent that knows the edge of its scope and hands off cleanly - a pattern worth designing deliberately rather than discovering in production. We covered the mechanics in structuring the AI-to-human handoff.
The single-scenario calls are worth more than dealers think
Here is the part that gets skipped. The bounded, unglamorous, single-domain calls are not a consolation prize. They sit on top of the largest documented revenue leak in the store.
Cox Automotive's 2026 research found that 80% of new-vehicle buyers say they would likely service at the selling dealership, but only 25% are introduced to the service department during the purchase and only 23% have a first appointment scheduled for them. The consequence compounds: buyers who return for service are 74% likely to repurchase from that dealer, against 44% for those who do not (Cox Automotive, April 2026).
The same study shows where that gap has led. Average dealer service and parts revenue reached $9.23 million, up 33% since 2018, while dealer share of total service visits fell to 29%. Among vehicles less than two years old - the vehicles dealers should own outright - share of service visits dropped from 68% to 55% over the same period.
Booking a first service appointment is a single-scenario task. So is confirming hours, quoting a menu-priced maintenance item, checking repair-order status, or recovering a declined service. These are the tasks that benchmark near the top of the range rather than the bottom, and they are the tasks a dealership's phones drop most often, because they arrive during the same hours the team is busiest. The cost of those missed calls is the easiest number in the store to recover.
Consumers are already meeting dealers halfway. Cox found 16% of consumers used an AI tool to research their last service provider, with price comparison and identifying needed services the leading use cases. But only 46% say they are likely to trust an AI recommendation for vehicle servicing. Trust is earned on execution, and execution is far easier to guarantee inside a narrow scope.
The market is about to make scoping harder, not easier
On July 23, 2026, HubSpot launched Agent Hub and Agent Builder in public beta for all Professional and Enterprise customers - a single place to build, monitor, and manage AI agents that share customer context, with a no-code canvas for assembling custom agents (HubSpot; CMSWire). Automotive CRM vendors will ship comparable tooling, and several already have pieces of it.
The strategic effect is worth naming. When building an agent takes an afternoon, the constraint stops being whether you can build it and becomes whether that agent should own the task. Stores that have not thought about scope will accumulate agents the way they once accumulated point solutions - several of them talking to the same customer, none of them accountable for an outcome. Shared context helps, and it is a real improvement over disconnected bots, but it does not by itself decide which conversations an agent should be allowed to finish.
A scoping test to run before you sign anything
Five questions, applied to each call type you are considering automating:
- How many systems must the agent touch to finish? One is a strong candidate. Three or more is a transfer.
- Can the agent be wrong in a way that costs money? Quoting a payment, a trade number, or an approval belongs to a person.
- Is there a defined finish line? Appointment booked, status confirmed, callback scheduled. If done is fuzzy, the scope is fuzzy.
- What happens when information is missing? Confirm the agent asks rather than assumes. This is the benchmark's documented failure mode, and it is testable in a pilot.
- How is the handoff logged? A transfer that loses context is worse than a voicemail, and it is where most of the customer-experience damage in these deployments actually occurs.
Then measure the narrow thing. Not AI adoption. Answer rate on service calls during peak hours. First appointments booked at delivery, against that 23% baseline. Declined-service recapture rate. Those are numbers a general manager can defend in a 20 Group.
The measured take
The 32.5% figure will get quoted by skeptics as proof that AI agents are not ready, and waved off by vendors as a benchmark artifact that does not reflect production systems. Both readings miss it. The finding is not a verdict on the technology. It is a map of where the technology currently works - and it points straight at the part of the dealership with the clearest documented revenue gap and the least complex conversations.
Dealers who scope narrowly, measure one metric, and expand only after the first task holds up will look conservative this year and well ahead in two. Dealers who buy the demo and hand over the switchboard will learn the difference between 32.5% and 62% on live customers.
If you are working out which of your calls belong in which bucket, that is the conversation we have with dealers most weeks. Carbuki builds AI voice agents scoped to the calls dealerships actually drop - worth a look as you plan the second half. Related reading: AI service scheduling in practice.
Sources
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications - Meituan LongCat, ICLR 2026. Repository and paper
- What Dealers Are Leaving in the Service Lane - Cox Automotive, April 2026, drawing on the 2025 Fixed Ops and Ownership Study, 2025 Service Industry Study, and 2025 Car Buyer Journey Study. Link
- Q2 New-Vehicle Purchases with Negative Equity Trade-Ins Hit Record Monthly Payments and Interest Costs - Edmunds via GlobeNewswire, July 16, 2026. Link
- Meet Agent Hub and Agent Builder - HubSpot, July 2026. Link
- HubSpot Debuts Agent Hub to Unify AI Agents - CMSWire, July 2026. Link
Carbuki builds AI voice agents for retail automotive — answering sales and service calls, following up on leads, and booking appointments 24/7 in multiple languages.
See how it works →