July 2026

The Prediction Is Already Becoming Testable

Eighteen days after I predicted Anthropic would be the first major frontier-model provider to fail, the mechanism behind that prediction has started to become testable.

The AI Stack
Eighteen days after I predicted Anthropic would be the first major frontier-model provider to fail, the mechanism behind that prediction has started to become testable.

On June 20, 2026, I published a prediction that Anthropic would be the first major frontier-model provider to fail. I defined failure as folding, being acquired under pressure, or losing enough market share that Claude becomes a niche product rather than an independent frontier competitor.

Checkpoint published July 9, 2026.

I also gave the prediction a deadline. By June 2028, I expect Anthropic’s share of the business AI coding market to fall from approximately 50% to approximately 20%.

Eighteen days later, SpaceX launched Grok 4.5.

I am not declaring victory. Anthropic has not lost the predicted market share. Claude remains the standard against which coding models are measured, and one product launch does not establish a market trend.

The prediction has not come true. The mechanism behind it has started to become testable.

What Arrived

Grok 4.5 is aimed directly at coding, agentic work, and the enterprise market. SpaceXAI says it trained the model across tens of thousands of NVIDIA GB300 GPUs using datasets focused on coding, science, engineering, and mathematics. Cursor participated in the training, and the model launched inside Cursor, Grok Build, and the SpaceXAI API.

The benchmark results are mixed, which is exactly why this should be treated as evidence rather than a verdict. Grok 4.5 beat Opus 4.8 on SpaceXAI’s reported DeepSWE 1.0, SWE Marathon, and Terminal Bench 2.1 results. Opus 4.8 remained ahead on DeepSWE 1.1 and SWE-Bench Pro. Different harnesses and evaluations produce different rankings, but the overall picture is clear enough: Grok has entered the competitive range for serious coding work.

That matters because my prediction never required SpaceX to build a model that defeats Claude on every benchmark. The competitive bar is acceptable parity at a better price.

Grok 4.5 is priced at $2 per million input tokens and $6 per million output tokens. Claude Opus 4.8 is priced at $5 and $25. The output price difference is especially important for agentic coding, where a model may produce and revise large amounts of code while working through a task.

SpaceXAI also claims Grok 4.5 runs at 80 tokens per second and used 15,954 output tokens per resolved SWE-Bench Pro task, compared with 67,020 for Opus 4.8. That is 4.2 times fewer output tokens. Those are SpaceXAI’s measurements, not independent proof of a production cost advantage, but they give us a concrete claim to test.

What This Proves

One of the original essay’s leading indicators was explicit: SpaceX would deliver a competitive coding model and price it below Claude. That has now happened.

Another indicator was that Cursor would shift more usage toward Grok, Composer, or other non-Anthropic models. Grok 4.5 launching inside Cursor does not prove that usage has shifted, but the distribution path now exists. Developers do not need to adopt a new editor, move their repositories, or redesign their workflows. Cursor can put the alternative directly beside Claude.

The third indicator was that SpaceX’s low-level training and inference work would produce a meaningful cost or speed advantage. The new pricing, 80-token-per-second serving speed, and token-efficiency claims are early evidence in that direction. They do not yet prove that the C stack caused the advantage or that the economics will hold under sustained production use.

We have, however, moved from hypothetical mechanisms to measurable claims.

What This Does Not Prove

It does not prove that Grok 4.5 is better than Claude. The benchmark results do not support that conclusion. It does not prove that developers will prefer Grok after using it on real codebases. It does not prove that Cursor will make Grok the default or steer meaningful usage away from Anthropic.

It also does not prove that SpaceX’s lower price is sustainable. The company may be subsidizing adoption, using available capacity aggressively, or accepting lower margins to gain market share. That would still create competitive pressure, but it would be a pricing strategy rather than proof of a permanent infrastructure advantage.

Most importantly, Anthropic’s market share has not yet moved in the way I predicted. The June 2028 prediction is about sustained commercial displacement, not launch-day benchmarks.

What Comes Next

The next evidence will come from behavior rather than announcements. Does Grok 4.5 perform well enough on real engineering work that teams keep using it after the free trial? Does Cursor change its defaults, pricing, or routing to favor SpaceX models? Does Anthropic respond by lowering prices, increasing subscription limits, or spending more to secure compute?

Then comes the number that matters: market share.

If Anthropic continues to control roughly half of the business AI coding market, this release is simply another capable model entering a crowded field. If Cursor usage shifts, Grok holds up in production, and Anthropic’s share begins a sustained decline, the first shoe will have dropped.

When that happens, I will write the next update.

Receipts

← All writing