July 2026

What AI Companies Can Learn From the Auto Industry

Frontier labs do not need one model for every job. They need a product ladder that gives customers Toyota economics, Porsche capability, and Koenigsegg performance when the work earns it.

The AI Stack
Frontier labs do not need one model for every job. They need a product ladder that gives customers Toyota economics, Porsche capability, and Koenigsegg performance when the work earns it.

I really want a Koenigsegg. If you’re not familiar, Koenigsegg makes hypercars and I think they’re amazing. They combine incredible technological innovation, great styling, and mind-bending performance. They are rare, technically extreme, and built for buyers who want performance that most drivers will never use.

I’d be very happy with a Porsche. Porsche sells performance that still fits daily life. The Panamera is my favorite because I’m a fan of go-fast wagons. (My current ride is a Lincoln Aviator.)

What does this have to do with AI? Compare Koenigsegg, Porsche, and Toyota. Koenigsegg sells hypercars. Porsche sells performance autos. Toyota sells the competent, dependable cars most people can use every day.

AI companies are treating the Koenigsegg as the default car, then subsidizing it. The auto industry learned a long time ago that markets do not work that way.

Frontier labs have a product strategy problem disguised as a token-pricing problem. They keep releasing smarter models. Those models use more reasoning, call more tools, carry more context, and run longer agentic workflows. More tokens get used. The work gets better. The bill gets bigger.

Premium products are normal in every market. A Porsche costs more than a Toyota because it does something a Toyota does not do. A Koenigsegg costs more than a Porsche because it does something even fewer buyers need.

The problem begins when the company expects customers to drive a Koenigsegg to the grocery store.

Context, July 2026: API prices, model names, and benchmark standings are moving quickly. This essay examines how a lab keeps customers as capability and cost move at different speeds.

The market already has a ladder

Brad Gerstner made the case on the July 10, 2026 All-In episode (280): every market has room for premium, mid-tier, and commodity products. He also argued that there is no evidence today that the intelligence gap between frontier and commodity models has collapsed. And he’s right about the current market.

The frontier labs still own work where a missed detail, a bad decision, or a failed agent costs more than the model bill. That is complex coding, long-running research, high-stakes analysis, and work that requires real judgment across a messy system.

Cheap models can take enormous volume without taking that spend. Vercel’s May production data showed DeepSeek with 49% of coding-agent token volume but only 4% of cost. Anthropic had 28% of tokens and 70% of cost in that segment. Cheap models were handling quantity. Frontier models were still getting paid for the expensive work.

The tiered market is forming in public.

CarAI roleWhat the buyer is paying for
ToyotaCommodity and mainstream modelPredictable cost, acceptable quality, and broad availability
PorschePremium professional modelBetter judgment, reliability, and performance on hard work
KoenigseggMaximum-effort frontier modelScarce capability for rare work where cost is secondary

The mistake is assuming every task belongs in the Koenigsegg row. Anthropic already has the beginning of this ladder: Haiku, Sonnet, and Opus, with Fable/Mythos as a fourth tier. That does not weaken the argument. It makes the next questions more important: which tier is the default, how quickly does capability move down the ladder, and does the product help a customer choose based on the work rather than the model’s prestige?

A catalog of models is not yet a product strategy. The strategy appears in the defaults, the routing, the permissions, and the bill.

Token cost pressure turns model choice into procurement

The price per token is only part of the cost. Agentic systems compound it. An agent reads files, calls tools, retries failures, adds history to context, reviews its own work, and sometimes launches subagents. A prompt that looked cheap as a chat interaction becomes expensive as a persistent worker.

That cost lands in an organization before the organization has a clean way to measure value. One employee can use a premium coding model for an afternoon and consume a large part of a shared team allocation. A business user can select the most capable model because the interface makes it available, not because the task needs it.

Then finance asks the question that every AI vendor eventually has to answer: what did we get for the spend?

The answer cannot be “more tokens.” It cannot even be “the strongest model.” It has to be an accepted result, delivered faster or more reliably than the alternative.

The price of intelligence must reflect the value of the workload, not the prestige of the model.

Nobody pays Porsche prices for a Toyota. Nobody pays Koenigsegg prices for a Porsche.

The Hypercar Trap

A frontier lab enters the Hypercar Trap when customers stop using it by default. Here is the sequence:

Flagship models become more capable and more expensive
→ routine workloads fail the ROI test
→ teams route routine work to cheap APIs, open models, or local inference
→ premium models become exception paths
→ the lab keeps prestige but loses the volume workload

The lab can still own the best model, win the leaderboard, and command a premium for the hardest work. It becomes a narrow supplier when it does not offer an affordable path for the ordinary work around that hardest work.

This is why a lower-cost competitor matters even when it is not the best model. Grok 4.5 launched at $2 per million input tokens and $6 per million output tokens. Claude Opus 4.8 is priced at $5 and $25. GPT-5.5 is priced at $5 and $30. Those are not small differences when an agent is generating and revising thousands of output tokens.

The competitive question is not whether Grok can beat every frontier model on every benchmark. The competitive question is whether it can do enough of the customer’s real work at a cost that makes the premium default look irresponsible.

Capability has to cascade

Toyota did not keep the Camry frozen while putting every improvement into a limited-production supercar. Safety, reliability, fuel efficiency, manufacturing quality, and useful technology moved down the lineup over time.

That is the model frontier labs need. The flagship should establish the frontier, prove new capabilities, and serve the rare tasks where a premium is earned. Then yesterday’s frontier capability needs to move down.

Anthropic is pursuing the same shape with Haiku, Sonnet, Opus, and Fable/Mythos. OpenAI’s GPT-5.6 family makes the structure especially explicit in its naming. Sol is the flagship. Terra is the lower-cost professional tier. Luna is the fastest and most affordable tier. OpenAI says those names are durable capability tiers that can advance on their own cadence.

That is the right product shape. A lower-priced model is not a permanent consolation prize. It is the place where improvements become normal.

The product strategy is a capability cascade:

  1. Build the Koenigsegg to extend the frontier.
  2. Move its proven capabilities into the Porsche tier.
  3. Move the now-common capabilities into the Toyota tier.
  4. Make the Toyota tier the default for ordinary work.
  5. Make escalation to Porsche or Koenigsegg intentional, visible, and economically justified.

The lab that does this well cannibalizes itself before a cheaper competitor does it for them.

The router is the dealership

Customers should not need an AI PhD to decide which model to use for every request. The product should know the available models, the user’s permissions, the cost budget, the data classification, and the task’s difficulty. It should send routine work to the Toyota. It should escalate difficult work to the Porsche. It should make the Koenigsegg available when the user has a real reason to use it.

Anthropic already demonstrates that this kind of routing is possible. Fable can route work to Opus when security restrictions require it. The mechanism is useful, but the framing matters. When routing is imposed as a provider restriction, customers experience it as a loss of choice.

My suggestion: give the customer an opt-in setting called efficiency routing. The product can then say, “This task can run on a lower-cost model and save your team money,” while keeping a premium route available when the work earns it. That is the product acting as the customer’s advocate.

That router is not merely a cost-control feature. It is how a lab protects its place in the customer’s architecture. If the lab does not provide it, the customer will build it. Or the customer will buy it from an AI gateway, a cloud provider, an IDE vendor, or a managed open-model platform.

Once that happens, the lab is no longer the default intelligence layer. It is one supplier inside somebody else’s routing policy.

Premium intelligence still has a future

Frontier intelligence has a future when it earns its premium. Gerstner’s market structure and the token-cost pressure can both be true. There can be a durable premium tier. There can be a growing commodity tier. There can be an expanding overall market for intelligence.

But premium providers cannot confuse a bigger market with permission to price every workload like a premium workload. The labs that win will make customers feel the difference between a Toyota, a Porsche, and a Koenigsegg. More importantly, they will make it easy to pay for the right one.

Receipts

  • All-In market-tier argument: Brad Gerstner’s July 11, 2026 appearance on All-In argued that every market supports premium, mid-tier, and commodity products, while noting that the frontier-to-commodity intelligence gap has not yet collapsed. Episode listing and timestamps.
  • Observed routing split: Vercel’s May 2026 AI Gateway data reported DeepSeek at 49% of coding-agent token volume and 4% of cost, while Anthropic had 28% of volume and 70% of cost. This is gateway-specific production data, not the entire market. Vercel Production Index.
  • Grok 4.5 pricing: SpaceXAI lists Grok 4.5 at $2 per million input tokens and $6 per million output tokens. SpaceXAI.
  • Claude Opus 4.8 pricing: Anthropic lists standard Opus 4.8 pricing at $5 per million input tokens and $25 per million output tokens. Anthropic.
  • GPT-5.5 pricing: OpenAI lists GPT-5.5 at $5 per million input tokens and $30 per million output tokens. OpenAI.
  • OpenAI’s tiered family: OpenAI describes Sol, Terra, and Luna as durable capability tiers, priced at $5/$30, $2.50/$15, and $1/$6 per million input/output tokens, respectively. GPT-5.6 announcement.
  • Routing observation: In the Claude desktop app, Fable routes some work to Opus when security restrictions apply. This is product behavior observed by the author, not a claim that Anthropic optimizes routing for workload value.
← All writing