Carbuki Insights
AI Got Cheaper Again This Month. Your Dealership's AI Bill Barely Moved - Here's Why.
Provider list prices for input tokens, mid-July 2026. The output-token spread is wider still. Sources: Techzine; explainX (2026).
The cost of raw AI keeps falling. Your monthly invoice has not noticed.
This week brings another wave of cheaper, more capable AI. DeepSeek is slated to push its V4 model to a stable release, and Moonshot AI has said it will publish open weights for its Kimi K3 model within days - the tail end of what industry trackers are calling one of the busiest stretches of open-model releases on record (BuildFastWithAI, July 20, 2026). It follows a sharper event two weeks earlier: on July 8 and 9, xAI took Grok 4.5 public and OpenAI opened its GPT-5.6 family to general availability inside roughly 48 hours, touching off a visible price fight over what top-tier reasoning should cost (Techzine, July 9, 2026).
For a dealer, the practical question underneath the headlines is fair: if the AI itself is getting this much cheaper, does that mean the AI on my phones, in my BDC, and on my website is about to get cheaper too? The honest answer is a little, but far less than the headlines imply - because the price that is collapsing was never the big number on your bill.
Myth: Now that AI models are getting dramatically cheaper, the cost of running AI at my store will fall just as fast.
Data: The model is one line in the bill, and it was already small. The cost to query a model at GPT-3.5 quality fell about 280-fold - from $20.00 to $0.07 per million tokens - between late 2022 and late 2024 (Stanford HAI, AI Index 2025). Yet only about 15% of franchised dealers have actually embedded AI into their workflows, while roughly 60% are still testing (Cox Automotive, 2025). The token price was never the thing holding adoption back.
The number that fell
Start with the raw trend, because it is genuinely dramatic. Stanford's AI Index found that the price to run a query at the quality of GPT-3.5 dropped from $20.00 per million tokens in late 2022 to about $0.07 per million tokens by October 2024 - a reduction of more than 280 times in roughly a year and a half (Stanford HAI, AI Index 2025). Depending on the task, independent trackers put the annual decline anywhere from 9 times to 900 times (Epoch AI, 2025). The July 2026 launches are the latest chapter of the same story: competitive pressure pushing capability up and price down at once.
The mid-July price sheet makes the fight concrete.
| Model (provider) | Input, per 1M tokens | Output, per 1M tokens | Public launch |
|---|---|---|---|
| GPT-5.6 Luna (OpenAI) | $1.00 | $6.00 | Jul 9, 2026 |
| Grok 4.5 (xAI) | $2.00 | $6.00 | ~Jul 8, 2026 |
| GPT-5.6 Terra (OpenAI) | $2.50 | $15.00 | Jul 9, 2026 |
| GPT-5.6 Sol (OpenAI) | $5.00 | $30.00 | Jul 9, 2026 |
Prices as published by the providers and reported in mid-July 2026 (Techzine; explainX, 2026). Numbers like these are why a single AI-handled phone conversation costs, in raw model terms, close to nothing. A few minutes of dialogue might exchange several thousand tokens; even on generous assumptions, that is a cent or two of model cost per call at today's prices. When that penny gets cheaper, you save a fraction of a penny.
What actually sits inside the cost of an AI-handled call
If the model is a rounding error, where does the money in an AI phone or chat tool actually go? Roughly five layers, only one of which the price war touches in a meaningful way:
- The model (inference). The tokens. Already pennies, and getting cheaper. This is the layer the July launches move.
- Speech. Turning caller audio into text and text back into natural speech, in real time, without awkward pauses. Priced by the minute, and largely independent of model token prices.
- Telephony. Carrier minutes, phone numbers, and call routing. A per-minute cost set by the phone network, not by any AI lab.
- The platform. The software that ties those pieces to your CRM and DMS, logs the lead, books into your scheduler, follows your rules, and does not embarrass you in front of a customer. This is engineering and integration - and it is where most of a vendor's invoice actually comes from.
- The process around it. The human work that turns a handled call into a sold car or a booked repair order: the follow-up, the confirmation, the manager who reviews what the AI did. It never appears on the vendor's line items, yet it is the real determinant of whether any of this pays.
A cheaper model lowers the first item and leaves the other four roughly where they were. That is why "AI got 280 times cheaper" and "our AI vendor's price barely moved" can both be true at the same time.
Where the money actually leaks
None of this makes the AI question small. It means the number worth watching is not the token price - it is the demand you are already paying to generate and then losing at the door. Franchised dealers spend heavily to make the phone ring; advertising alone runs on the order of several hundred dollars per new vehicle sold (NADA Data, 2025, via industry reporting). What happens to those calls is the leak. Independent and vendor call studies converge on the same uncomfortable range: a large share of inbound dealership calls - commonly cited between one in five and one in three - go unanswered, and a majority of leads now arrive outside core showroom hours (industry call-tracking analyses, 2026).
That gap is what AI is actually being bought to close, and it explains the Cox Automotive finding that dealers do not care about AI for its own sake - they care about outcomes they can measure, such as more cars sold, lower inventory costs, and higher gross profit (Cox Automotive, 2025). In that same study, 81% of dealer leaders said AI is here to stay and 63% called investing now critical for long-term success - yet only about 15% have embedded it into operations (Cox Automotive, 2025). What stands between testing and embedding is not the cost of a token. It is trust, integration, and process.
What cheaper models change - and what they do not
Falling model prices are real leverage, just not on the invoice line most owners expect. Three effects are worth taking seriously:
- More capability at the same price. The clearest benefit is not a smaller bill; it is a better agent for the same money. Cheaper, stronger reasoning lets a phone or chat agent handle messier, multi-step conversations - a trade-in question that becomes a service question that becomes an appointment - without the cost spiking.
- A lower bar to test. When the underlying model is cheap, vendors can offer trials and smaller-store pricing that were harder to justify before. That helps the 60% still testing, and it is part of why some manufacturers have started folding these tools into co-op programs.
- Less reason to tolerate a bad experience. If capability is rising while cost falls, "the AI sounded robotic and dropped the call" is harder to excuse than it was a year ago. The floor for what counts as acceptable is moving up.
What cheaper models do not change: whether the tool is wired into your CRM and DMS, whether it logs every lead, whether it books into the right scheduler, and whether your team acts on what it captures. Those variables decide ROI, and they are stubbornly human and operational.
How to read this as a buyer
The takeaway is not "wait for AI to get cheaper." It already did, and the savings mostly accrue to model providers and vendors, not to the buyer's bottom line. The takeaway is to judge AI at your store the way you judge a salesperson: on outcomes, not on input costs. A few questions cut through the noise:
- Does the tool connect to the systems you already run, or does it create a second place to check?
- Can it show, in your own reporting, the calls answered, appointments set, and leads logged that would otherwise have been missed?
- When it fails, how does it hand off to a human, and can you see what it did?
- Is the trial structured around a measurable outcome over 60 to 90 days, rather than a feature demo?
A store that answers those well will do better with a mid-priced tool than a store that buys the cheapest model and changes nothing about its process. For a fuller version of that argument, see our earlier piece on the execution gap between AI that answers and AI that finishes the job.
The falling cost of intelligence is good news, and it is not slowing down. But the number that decides whether AI pays off at your store is the one you have always watched: how much of the demand already calling you gets answered, booked, and logged. That is the problem Carbuki is built for - AI voice agents that answer every call, book the appointment, and make sure no lead goes missing. If you want to test the idea against your own phone logs, it is a conversation worth having.
Sources
- Stanford HAI, "The 2025 AI Index Report - Research and Development" (inference cost decline of roughly 280x): https://hai.stanford.edu/ai-index/2025-ai-index-report/research-and-development
- Epoch AI, "LLM inference prices have fallen rapidly but unequally across tasks" (2025): https://epoch.ai/data-insights/llm-inference-price-trends
- Techzine, "GPT-5.6 now widely available: Sol, Terra, and Luna launched," July 9, 2026: https://www.techzine.eu/news/applications/142797/gpt-5-6-now-widely-available-sol-terra-and-luna-launched/
- explainX, "Grok 4.5 Public Launch: Benchmarks and Pricing," July 2026: https://explainx.ai/blog/grok-4-5-public-launch-spacexai-july-2026
- BuildFastWithAI, "AI News Today, July 20, 2026" (late-July open-weight release wave; DeepSeek V4, Moonshot Kimi K3): https://www.buildfastwithai.com/blogs/ai-news-today-july-20-2026-16-biggest-stories
- Cox Automotive, "Automotive Dealers Are Ready for AI to Deliver Outcomes and Skip the Hype, According to New Cox Automotive Study," October 2025: https://www.coxautoinc.com/insights/automotive-dealers-are-ready-for-ai-to-deliver-outcomes-and-skip-the-hype-according-to-new-cox-automotive-study/
- NADA, "2025 Annual Financial Profile of America's Franchised New-Car Dealerships" (dealership advertising and financial benchmarks): https://www.nada.org/nada/nada-data
Carbuki builds AI voice agents for retail automotive — answering sales and service calls, following up on leads, and booking appointments 24/7 in multiple languages.
See how it works →