Carbuki Insights

Your AI Vendor and Its Competitor May Run the Same Model. The Difference Is Everything Around It.

August 15, 2026

Dealers working with an external AI partner report better outcomes (2026)
Say they are using AIoptimally66%Report sales and revenuegrowth from AI36%Express high confidence inAI outputs30%

Share of dealers who work with an external AI partner. Without a partner, the same three figures are 46%, 21% and 8%. This is correlation, not proof of cause - but the gap sits in implementation, not in model choice. Source: Cox Automotive AI in Auto Retail Tracker, combined Q1-Q2 2026 base (504 dealers surveyed in Q1, 483 in Q2).

The week the model got cheaper, faster, and less important

Three announcements landed inside 48 hours this week, and together they say something useful to anyone at a dealership currently evaluating an AI phone or BDC tool.

On August 13, Google released Gemini 3.7 Flash with introductory pricing of $0.75 per million input tokens and $3.75 per million output tokens - half the cost of the previous Flash version (Reuters, August 13, 2026). The same day, OpenAI previewed Ultrafast, a service tier running GPT-5.6 Sol at up to 750 output tokens per second, as much as 14 times faster than standard processing, on hardware from Cerebras (OpenAI and Cerebras, August 13, 2026). Also that day, the enterprise AI company Writer published research arguing that the software wrapped around a model - what engineers call the harness, or the orchestration layer - drives cost and speed at least as much as the model itself does (TechCrunch, August 13, 2026).

DeepSeek, meanwhile, went the other direction, repricing some workloads upward by as much as 1,100% with its new V4 Pro flagship (Reuters, August 2026).

Four data points, one conclusion for a dealer: the model inside your AI phone agent is becoming a commodity input with a volatile price. It is not the thing you are buying. What you are buying is everything built around it.

Myth: Choose the vendor with the best AI model and you will get the best results.

Data: Writer reports that rebuilding only its orchestration layer - holding the model constant - cut cost per task by 41% and completed tasks 44% faster across every model it tested, including third-party models from Anthropic and OpenAI. The model did not change. The outcome did.

Source: Writer, reported by TechCrunch and VentureBeat, August 13, 2026. Vendor-reported figures; treat them as directional rather than audited.

What a harness is, in dealership terms

A large language model is a component. On its own, it produces text. Everything that turns that text into a handled phone call is separate software: the code that decides what information to hand the model, which tools it is allowed to call, how long it waits, what happens when a tool fails, how much of the earlier conversation it carries forward, when it stops and passes the call to a person, and what it writes back to your systems afterward.

That is the harness. At a store, it is the difference between:

  • A model that can say the Tacoma is available, and a system that checked your live inventory feed before saying it.
  • A model that produces a tidy summary, and a system that writes a structured record into your CRM with a next step attached.
  • A model that answers, and a system that offers two real times, books one on the real calendar, and sends the confirmation.
  • A model that keeps talking, and a system that recognizes an upset customer or a payoff question and hands the call off cleanly.

None of those behaviors come from the model. They come from the engineering around it - the same layer Writer measured, and the layer no two vendors build the same way even when they license identical underlying models.

The price and speed news still matters

Commodity inputs have economics, and this week they moved in a direction dealers benefit from.

Development (mid-August 2026)What changedSource
Google Gemini 3.7 FlashIntroductory pricing of $0.75 and $3.75 per million input and output tokens - half the prior Flash costReuters
OpenAI Ultrafast (GPT-5.6 Sol)Up to 750 output tokens per second, as much as 14x standard speed, limited previewOpenAI and Cerebras
OpenAI and Anthropic price cutsPrices paid for leading U.S. models have fallen materially since mid-JulyFinancial Times
DeepSeek V4 ProSome API workloads repriced upward by as much as 1,100%Reuters and Caixin

Two things follow. First, falling per-token cost is why always-on phone coverage keeps getting easier to justify. The marginal cost of answering the 11 p.m. service call is heading down, not up. We walked through how that cost stack works in the AI price war and your dealership cost stack.

Second, latency has become a product feature that vendors now buy separately - and it matters more on a phone call than almost anywhere else. In a chat window, a two-second pause is invisible. On a call, it is the pause that makes a caller start talking over the agent. Cerebras named voice and customer support explicitly among the target workloads for the faster tier. If you have tested a voice agent that felt stilted, you were often feeling infrastructure rather than intelligence, a point we made in full-duplex voice AI and the dealership phone.

But notice what neither development does: neither one makes a badly integrated tool work. A faster model books an appointment faster only if the software around it can book appointments at all.

The auto-retail data points the same way

The clearest evidence that the wrapper matters more than the model comes from the industry own numbers rather than from a lab.

Cox Automotive released its AI in Auto Retail Tracker on August 11, 2026. It found 82% of dealers using AI, 69% expecting AI to drive sales and revenue growth, and only 22% of AI users reporting they have actually seen that growth. About one in three either are not measuring the impact or have no clarity on how they are measuring it.

The more interesting split sits further down the release:

Outcome reportedWith an external AI partnerWithout
Say they are using AI optimally66%46%
Report sales and revenue growth from AI36%21%
Express high confidence in AI outputs30%8%

Source: Cox Automotive AI in Auto Retail Tracker, combined Q1-Q2 2026 base (504 dealers surveyed in Q1, 483 in Q2).

Read that carefully, because it is easy to over-claim. This is correlation, not proof of cause. Stores that engage a partner may already be better run, better capitalized, and better at measurement, and Cox does not argue otherwise.

Still, the direction lines up with what the engineering research says. Every dealer in that sample can reach roughly the same frontier models. The models are not what separates the 36% from the 21%. Implementation is: integration into the CRM and DMS, scoping the job narrowly, defining the handoff, and measuring the result. We covered the full tracker in the widening gap between AI-powered shoppers and dealers.

Six questions that test the layer around the model

If the model is the commodity, your evaluation should barely mention it. These six questions test the part that actually varies between vendors.

  1. What does it read before it speaks? Ask to hear a call where the agent confirmed live inventory or an open service bay. A system that cannot read your data is guessing politely.
  2. What does it write, and where? A transcript emailed to the desk is not a CRM record. Ask which fields it creates or updates, and look at a real example from another store.
  3. What happens when a tool fails? Scheduler down, DMS timing out, VIN not found. A good harness degrades into a clean handoff. A weak one improvises, which on a live call means inventing something.
  4. How does it decide to stop? Payoff amounts, negative equity, an angry customer, anything in F&I. Written escalation rules beat verbal assurances - see where voice AI should hand off to a human.
  5. How is latency measured, and what happens under load? Ask for time-to-first-word and behavior at your Saturday peak, not a demo at 10 a.m. on a Tuesday.
  6. Who is accountable when it gets something wrong? Model providers disclaim; your vendor should not. Consent, disclosure and record-keeping remain yours - see AI disclosure rules and the dealership phone and who owns AI at the dealership.

Notice that none of those questions is which model do you use. Right now, that answer has a shelf life of about three weeks.

The measured takeaway

The honest summary of this week is not that AI got better. It is that the part of AI everyone talks about - the model - got cheaper, faster and more interchangeable, while the part almost nobody markets - the orchestration around it - was shown to swing cost and speed by 40% or more on its own.

For a dealer, that is good news and a warning at once. Good news, because the input costs behind always-on phone coverage are falling, and the latency that made early voice agents feel robotic is being engineered out. A warning, because it means a slick demo running on a frontier model tells you almost nothing. Two vendors quoting you next month may be running the same weights. The one that lifts your appointment count will be the one that wired those weights into your inventory, your CRM and your service calendar - and that gave you a way to measure whether it worked.

So start with the metric, not the model: answer rate, after-hours capture, appointments set and kept, records logged. Cox data suggests the single biggest thing separating dealers who see revenue from AI from those who do not is whether anyone is measuring at all.


If you are evaluating an AI voice agent and want to pressure-test what sits around the model - the integrations, the handoff rules, and the numbers you would judge it by - that is the conversation the team at Carbuki has with dealers every day.

Sources

  • Reuters, Google unveils Gemini 3.7 Flash AI model for coding and agent workflows, August 13, 2026: reuters.com
  • OpenAI, Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed, August 13, 2026: openai.com
  • Cerebras, Accelerating GPT-5.6 Sol Ultrafast with OpenAI, August 13, 2026: cerebras.ai
  • TechCrunch, Writer introduces new AI model and upgraded harness to contain token costs, August 13, 2026: techcrunch.com
  • VentureBeat, Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges, August 2026: venturebeat.com
  • Cox Automotive, New Cox Automotive AI in Auto Retail Tracker Finds Growing Gap Between Dealers and AI-Powered Car Shoppers, August 11, 2026: coxautoinc.com
  • Tech Startups, Top Tech News Today, August 14, 2026 (summarizing Financial Times reporting on U.S. model price cuts and Reuters and Caixin reporting on DeepSeek V4 Pro pricing): techstartups.com

Carbuki builds AI voice agents for retail automotive — answering sales and service calls, following up on leads, and booking appointments 24/7 in multiple languages.

See how it works →
Share:XLinkedInFacebookRedditEmail

← All articles