Carbuki Insights
AI Vendors Now Charge More at Peak Hours. Your Service Drive Has Peak Hours Too.
Cache-miss output token pricing for V4-Flash under DeepSeek's published schedule, effective 16:00 UTC Aug. 16, 2026. Off-peak is half of peak; the same token costs about 4.7x more at peak than under the old flat rate. Source: InfoWorld, Aug. 13, 2026.
Sometime around noon Eastern on Sunday, the cost of running one of the industry's cheapest AI models started to depend on what time of day it is.
DeepSeek moved its V4 model family to peak and off-peak billing effective 16:00 UTC on August 16, with increases running from 50% to more than 1,100% depending on the model, the token type and the hour of use. Off-peak rates are half of peak rates. The company said the new schedule is intended to "allocate resources more reasonably" and to encourage developers to schedule work when capacity is less contended (Quartz, InfoWorld, 2026).
No US dealership is buying tokens from DeepSeek. That is not why this matters. It matters because of the assumption sitting underneath time-of-day pricing: that the work can wait for the cheap hour. Most of what happens at a dealership cannot. A customer standing in your service drive at 7:45 on a Tuesday is not a batch job.
Myth vs. data: AI costs only move in one direction, so waiting is the cheap strategy.
- In the same week, DeepSeek raised V4 prices by 50% to more than 1,100%, and Google cut Gemini 3.7 Flash prices (InfoWorld, Aug. 13-14, 2026).
- Info-Tech Research Group's Mark Tauschek noted Anthropic raised prices in April for the same reason: demand outrunning supply.
- Compute is now a two-way market priced by scarcity. That is worth knowing before signing a multi-year agreement priced off somebody else's cost curve.
What actually changed in the price schedule
The published rates, per InfoWorld's reporting on DeepSeek's schedule, per million tokens:
| Model and token type | Old flat rate | New off-peak | New peak |
|---|---|---|---|
| V4-Flash input (cache miss) | $0.14 | $0.22 | $0.44 |
| V4-Flash output | $0.28 | $0.66 | $1.32 |
| V4-Pro input (cache miss) | $0.435 | $0.66 | $1.32 |
| V4-Pro output | $0.87 | $1.98 | $3.96 |
Two details are more interesting than the headline percentage. First, Sanchit Vir Gogia of Greyhound Research told InfoWorld that 17 of every 24 hours stay at the half-price rate, so timing becomes an economic variable rather than a penalty. Second, the steepest increases (up to 1,100%) land on cached input tokens, which is exactly the mechanism that had kept DeepSeek's measured cost per task roughly 60% below a comparable OpenAI tier even after OpenAI cut GPT-5.6 Luna API prices by up to 80% in late July.
In other words: the sticker price moved, but the real variable is when and how the work is run. That is the same conclusion we reached looking at model interchangeability and orchestration - the model is rarely the thing that determines what you pay or what you get.
The assumption buried in the phrase off-peak
Time-of-day pricing works because a large share of enterprise AI work is genuinely elastic. Summarizing documents overnight, classifying a backlog, regenerating listings copy, running a coding agent at 3 a.m. - all of it can move to the cheap hours with no one noticing.
Retail automotive inverts that. Demand arrives in a spike, at the least convenient hour, from people who will dial the next store within minutes if nobody picks up. Your peak is not a scheduling preference. It is the appointment, the trade walk, the declined-work approval and the first service visit that either happens or does not.
So the useful exercise for a dealer is not "what does AI cost per minute." It is: which of my workloads are time-inelastic, and which are shiftable? Those two categories deserve completely different buying logic and completely different scoreboards.
Where the peak actually costs money: the service drive
Service is where the arithmetic is least forgiving, because the revenue is large, recurring and quietly contested. Cox Automotive's 2026 Fixed Operations and Ownership Study (April 2026, surveying 500 fixed-ops decision makers and 2,500 consumers) found average dealer service and parts revenue reached roughly $9.23 million in 2025, up 33% over eight years - while the dealer share of service visits fell from 33% to 29%. NADA's 2025 figures put industry service and parts sales above $164 billion across more than 276 million repair orders.
Revenue up, share down. That is a capture problem, not a demand problem. Cox counted nearly 299,000 auto mechanic businesses operating in the US, up 12% since 2018, plus mobile service as an entirely new category. And the price objection is largely perception: average consumer spend was $261 at a dealership versus $275 at general repair.
The study quantified four gaps that open up precisely when a store runs out of peak-hour capacity:
| Gap | What Cox measured | What it is worth |
|---|---|---|
| First appointment | 80% of new-vehicle buyers say they are likely to service at the selling dealership, but only about one in four had a first appointment scheduled at purchase - roughly a 50-point gap | Losing a service customer can represent more than $12,000 in potential lifetime service spend |
| Repurchase | Customers who return for service are 30 points more likely to repurchase; 88% say the service experience affects that decision | The next vehicle sale, not just the RO |
| Trade-in | Only 14% have ever been offered a trade value during a service visit, while 33% are highly interested | Service-lane acquisition; consumers begin weighing trade over repair around $3,195 |
| Transparency | Customers who received photos or video reported about $230 higher average repair orders; 49% say visuals make them more likely to approve work | Approval rate on recommended work |
Cox also found 16% of consumers used an AI website or tool during their most recent service journey - research, comparison, understanding a recommendation. Shoppers are arriving better informed on the service side too, which is the same pattern the AI in Auto Retail Tracker found on the sales side.
Sorting your own work: what can wait, and what cannot
This is the exercise worth doing on a whiteboard before the next vendor call.
Time-inelastic (must be handled live, at peak):
- Inbound service calls in the first and last operating hour
- Status calls on vehicles already in the shop
- Inbound sales calls on advertised or newly listed inventory
- After-hours, lunch-hour and weekend inbound
Shiftable (can and should be scheduled):
- Declined-service follow-up
- Recall and open-campaign outreach
- Maintenance-due and first-appointment reminders
- Unsold showroom and lease-end lists
- Equity and service-to-sales database mining
- Review and CSI requests
If AI vendors are increasingly pricing by the hour, shiftable work is where cost engineering belongs - batch it, schedule it, negotiate on volume. Time-inelastic work should be judged on capture rate instead. A call that gets dropped at 7:45 a.m. is not a cost saving; on Cox's own numbers it is a candidate for a five-figure lifetime loss.
This is also the honest limit on in-vehicle and OEM-side assistants, which we looked at here: they can route a request, but somebody at the store still has to answer it and put it on the schedule.
Questions worth asking before signing
- How is pricing structured - per minute, per call, per handled conversation, or per outcome? Does any part of it vary by time of day or concurrency?
- Is there a pass-through clause for upstream model price changes? April and August both showed those changes are not hypothetical.
- What happens at peak concurrency when 11 calls arrive at 7:45 a.m. - does it queue, degrade to a smaller model, or drop?
- Can it write, not just listen? Booking into the scheduler against real shop capacity is a different product from taking a message.
- How is capture measured, and can you see it by hour of day rather than as a monthly total?
That last question is the one most stores cannot answer today. Cox's tracker found 82% of dealers using AI, with about one in three either not measuring its impact or unclear how they measure it - and only 22% able to point to sales or revenue growth from it.
What to look at in 60 days
- Answer rate in the first and last operating hour, separated from the daily average
- Appointment-set rate on inbound service calls
- Share of new-vehicle deliveries that leave with a first service appointment booked (Cox's benchmark: about one in four)
- Declined-service follow-up contact and conversion rate
- Average repair order on visits where photos or video were sent versus not
Bottom line
The DeepSeek schedule is a small story with a durable lesson: AI capacity is now scarce enough that providers price it by the hour, and the cheap hour only helps if the work can wait. Dealership work mostly cannot. The stores that come out ahead will be the ones that separate the two categories deliberately - scheduling everything that can be scheduled, and making sure the hours that cannot be moved are actually covered.
If you want a clearer read on which of your hours are leaking calls, and which ones are worth covering first, that is the problem Carbuki works on.
Sources
- Quartz, DeepSeek is raising AI developer access prices by up to 1,100% starting Sunday, Aug. 13, 2026 - qz.com
- InfoWorld, DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity, Aug. 13, 2026 - infoworld.com
- InfoWorld, Google cuts Gemini 3.7 Flash prices as enterprise AI economics diverge, Aug. 14, 2026 - infoworld.com
- Cox Automotive, 2026 Fixed Operations and Ownership Study, April 9, 2026 - coxautoinc.com
- Cox Automotive, AI in Auto Retail Tracker, Aug. 11, 2026 - coxautoinc.com
- NADA, NADA Data: Annual Financial Profile of America's Franchised New-Car Dealerships - nada.org
Carbuki builds AI voice agents for retail automotive — answering sales and service calls, following up on leads, and booking appointments 24/7 in multiple languages.
See how it works →