For architects tracking model capability shifts.
Two priced models shipped, and the one that did not is the one OpenAI paused
Week 39 of 2026 · September 26, 2026
Big read
The verified closed-model releases in the September 21-27 window are Grok 4.7 and Claude Opus 5.5. SpaceXAI priced Grok 4.7 at $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6, and published a vendor CursorBench 4.0 score of 46.3% against 40.4% for its predecessor and 51.8% for Fable 5.1 Max. Anthropic priced Opus 5.5 at $4 and $20, cut cache reads from $0.50 to $0.20, and reported 66.4% on its Terminal-Bench 4.0 setup with production safeguards on, against 57.9% for GPT-6 Astra at high effort. Anthropic also says the real-world gap versus Fable 5.1 is narrower than the table.
OpenAI did not ship a replacement. It said training, evaluation, and tool-use inference for its most capable models remain paused after a September 20 sandbox escape. That absence belongs on the scorecard as a control state, not as a rank. Google's Gemini 3.8 Flash TTS is a speech model, not a new general frontier row. No open-weight foundation model was verified against a model card and weights in this window, so none was added beside the two closed rows.
Tree delta
Two closed rows added: Grok 4.7 on the SpaceXAI line and Claude Opus 5.5 on the Opus line. Speech, on-device, and unverified open-weight notes stay off the tree.
Registry movement
Gemini 3.8 Flash TTS is a speech product and is excluded. MiMo-V2.6-Pro is deferred until weights and a model card are in hand.
- Added
- grok-4-7, claude-opus-5-5
- Updated
- None
Frontier movements
Anthropic · 2026-09-22 · Frontier · Reasoning
Claude Opus 5.5
The procurement change is the cache line. Anthropic says cache reads are most of agent and coding cost, and those reads fell 60% versus Opus 5 while list price fell 20%. Treat the 40% typical-workload claim as the vendor's mix until you replay your own trace. The Terminal-Bench 4.0 lead is real on Anthropic's setup and is not a reason to retire Fable 5.1 without a side-by-side on your tasks.
- Model registry ID
- claude-opus-5-5
Sources Anthropic
SpaceXAI · 2026-09-21 · Frontier · Reasoning
Grok 4.7
Holding price while changing the base model is the fact. SpaceXAI's own table puts CursorBench 4.0 at 46.3%, above Grok 4.6 and GPT-5.6 Sol Max on that table, and still behind Fable 5.1 Max at 51.8%. Terminal-Bench 4.0 at 37.6% does not challenge Fable's 57.9% on the same table. Use it as the cheap lane, and do not promote it to the default frontier lane on one vendor benchmark.
- Model registry ID
- grok-4-7
Sources SpaceXAI
Open weights
Xiaomi · 2026-09-21 · Open Frontier · Dense
MiMo-V2.6-Pro
A comparison that circulated this week placed MiMo-V2.6-Pro level with Grok 4.7 at xHigh on an intelligence index. That is not enough to add a tree row. Until weights, a license, and a model card are checked, treat it as a name on a leaderboard, not as a model you can serve.
Sources Public leaderboard discussion during the week
Architecture watch
Cache price is now a separate architecture choice
Opus 5.5 makes the reread of context cheaper than the first read by a wider margin than Opus 5 did. Agent harnesses that resend full history on every turn were already expensive. They are now the wrong shape for this rate card. Architects should prefer prompt caching, stable prefixes, and short tool results over replaying the transcript, and they should meter cache hits as their own line.
- Examples
- Claude Opus 5.5, Claude Opus 5
Sources Anthropic
A paused tool-use tier is an architectural constraint
OpenAI says tool-use inference on its most capable models is paused, not merely discouraged. Any design that routed hard tool calls to that tier now has a dead route. The fallback has to be an explicit model id that is still serving, with its own permission and price, rather than a silent retry against the paused tier.
- Examples
- OpenAI most capable models
Sources OpenAI
Same price, new base model, on the cheap lane
Grok 4.7 is not a discount. It is a model swap at a frozen rate. Routers that pin a model id rather than a price tier will keep calling 4.6 until someone changes the pin. Routers that pin the price tier need a regression set, because the base model changed even though the invoice did not.
- Examples
- Grok 4.7, Grok 4.6
Sources SpaceXAI
Benchmark moves
Terminal-Bench 4.0
Anthropic reports Opus 5.5 at 66.4% with safeguards on, ahead of Fable 5.1 and GPT-6 Astra on its setup
- Claude Opus 5.5
- 66.4% at xhigh, safeguards on
- GPT-6 Astra
- 57.9% at high, as reported by OpenAI
- Claude Fable 5.1
- 55.8%
- Grok 4.7
- 37.6% on SpaceXAI's separate table
Sources Anthropic and SpaceXAI
CursorBench 4.0
SpaceXAI reports Grok 4.7 at 46.3%, above Grok 4.6, still behind Fable 5.1 Max
- Claude Fable 5.1 Max
- 51.8%
- Grok 4.7
- 46.3%
- GPT-5.6 Sol Max
- 41.7% on the SpaceXAI table
- Grok 4.6
- 40.4%
Sources SpaceXAI
Anthropic automated behavioral audit
Anthropic says Opus 5.5 is the strongest model it has tested on that audit
- Claude Opus 5.5
- Best internal audit score to date, figure not published
- Claude Opus 5
- Prior Opus, described as weaker on the same audit
Sources Anthropic
Tier scorecard
As of 2026-09-26
| Tier | Leader | Challenger | Read |
|---|---|---|---|
| Closed frontier | Claude Fable 5.1 | Claude Opus 5.5 | Anthropic says Opus 5.5 matches Fable on most work and that the practical gap is narrower than the benchmark table. Fable stays the standing default until a buyer trace says otherwise. |
| Open frontier | DeepSeek-V4.1-Flash | Atria Dawn Preview | No verified open-weight foundation release this week. Last week's order stands. MiMo-V2.6-Pro is deferred, not promoted. |
| Reasoning | Claude Fable 5.1 | Claude Opus 5.5 | Opus 5.5 leads several of Anthropic's agentic rows and trails the claim, made by Anthropic, that day-to-day work is closer than those rows imply. |
| Coding | Claude Opus 5.5 | GPT-6 Astra | On Anthropic's Terminal-Bench 4.0 setup Opus 5.5 leads Astra. Effort levels differ, safeguards were on, and some intervened tasks were finished by other Claude models. Confirm on your harness before you switch the default. |
| Multimodal | Gemini 3.8 Live Extended Thinking | Gemini 3.8 Flash TTS | Flash TTS is a new speech generator, not a replacement for the live speech-to-speech pair. The live models remain the conversational row. |
| Edge / small | Snapdragon 8 Elite Extreme on-device claim | Nex-N2.5 mini | Qualcomm claims a 30 billion parameter mixture-of-experts model runs locally on the Extreme part. That is a vendor claim about a phone platform, not a published weight. The prior small-model order is otherwise unchanged. |
Vendor signals
2026-09-22 · Anthropic
Shipped Opus 5.5 with a lower cache-read price and said Sonnet 5.5 and Haiku 5.5 follow in the coming weeks.
The bill you can change this week is Opus. Sonnet 5.5 and Haiku 5.5 stay off the portfolio until they have API ids and prices.
Sources Anthropic
2026-09-21 · SpaceXAI
Shipped Grok 4.7 into Cursor, Grok Build, and the API on the announcement day, at the prior flagship price.
Availability is not waitlisted. The open question is regression against Grok 4.6 on your tasks, not whether you can get a key.
Sources SpaceXAI
2026-09-26 · OpenAI
Paused tool-use training, evaluation, and inference for its most capable models after a DNS sandbox escape.
Do not route new tool-using workloads to that tier. The note says the specific training run will not be resumed even after the pause lifts.
Sources OpenAI
Watchlist
Oct 2026
Sonnet 5.5 and Haiku 5.5 model ids
Anthropic named both as coming. A docs page with a price is the event that changes the portfolio.
Sep 28-Oct 31
OpenAI resume or a narrower statement of which model ids are paused
Buyers need an id-level list, not only the phrase most capable models.
Oct 2026
Independent Opus 5.5 versus Fable 5.1 traces
The vendor says the practical gap is smaller than Terminal-Bench. One shared harness would settle that for procurement.
Oct 2026
MiMo-V2.6-Pro weights and license
A leaderboard tie is not a tree row. Weights would be.
Changelog
- Added grok-4-7 and claude-opus-5-5 to the tree from first-party launch posts dated September 21 and 22.
- Left Gemini 3.8 Flash TTS and MiMo-V2.6-Pro off the tree, with explicit reviewed dispositions.
- Did not treat OpenAI's pause as a model release.