For architects tracking model capability shifts.
Same list price, three different bills
Week 40 of 2026 · October 3, 2026
Big read
A 'model' is now a policy-wrapped system, and its price is a set of meters rather than a list rate: that is what this week's three mid-tier releases show, and it changes how you evaluate and how you budget. Start with the gates, because each lab shipped a different kind. Google's is an access gate: Gemini 4 Argon exists, scores on independent indices, and cannot be bought. Anthropic's is a routing gate: Sonnet 5.5 can hand a higher-risk cyber request to Sonnet 5 mid-session, visibly, through classifiers in the request path. OpenAI's is a classification gate: Sol inherits GPT-6 Astra's safeguards stack with a Critical rating in cybersecurity (OpenAI's highest internal capability tier, which triggers its strictest deployment controls), and the planned Astra successor was withheld with no numbers. The September 25 pause covers an unreleased most-capable tier; GPT-6 Astra itself still serves dots and the Ultrafast tier. An evaluation harness (the software that wraps the model, supplies its tools and runs the benchmark) that assumes a fixed model behind an endpoint will produce numbers that do not reproduce in production.
The meters are the second half. For a session that rereads a large context many times, the cache-read meter alone can make the Sol and Sonnet 5.5 bills diverge by 2x behind the same $2 / $10 list price; for a single long-document pass, Sol's long-context repricing does the opposite. Argon's $2 / $10 is a 50% introductory discount for at least one month, and TokenCost noted that its post-intro card is Opus 5.5's card line for line; DataCamp and TokenCost both published the meters-over-list-price read before this issue, and synthorai added the multiple the meter table does not show, that Sonnet 5.5 writes 1.7-3.1x the output tokens of Sol on comparable tasks. The full per-meter table, with the cache-read rates and the 272K threshold, is in the architecture watch below. Above the mid tier on the independent index, Opus 5.5 at $4 / $20 is the only buyable model that outscores Sonnet 5.5, and the question for most agent budgets is whether its two extra index points are worth doubling the bill; GPT-6 Astra at $10 / $50 is buyable and scores below Sonnet 5.5, which is the stronger point for a buyer.
One number to carry: Anthropic's own Terminal-Bench 4.0 harness puts Sonnet 5.5 at 70.6% and Artificial Analysis's independent run (its September 30 Argon article) at 64%, a 6.6-point gap that is the size of the vendor-versus-independent difference to expect in your own evaluation; the independent table, with the effort settings each model was run at, is in the benchmark moves.
What to do: rerun the agent cost model with cache-read rates, long-context thresholds and output tokens per task as inputs, not list price; benchmark Sonnet 5.5 against Sol on your own task set rather than theirs, with the real request mix including the requests the gates are designed to catch; and treat Argon as a Q4 evaluation item, not a Q4 deployment option. The capital and power consequences are in AI Stack Weekly; the permission-layer consequences for agents are in Agent Techniques.
Tree delta
Six rows added, one updated. Three closed frontier-adjacent models at $2 / $10 (Claude Sonnet 5.5, GPT-6.1 Sol, Gemini 4 Argon as gated), two open-weight MoE releases (mixture-of-experts, a model that activates only a few of its sub-networks per token; MiMo-V2.6-Pro now verified, Kolibri-1), and Holo4 under the Holo3.1 precedent (W23 placed H Company's computer-use family as a tree node because it ships in multiple sizes with quantized checkpoints, reduced-precision weights, for local inference). The decision-model fine-tunes are reviewed and deferred rather than placed: the tree's 3-of-6 rule creates a new branch only when at least three of six tests hold (durable architectural distinction, three or more independent vendors, changed infrastructure behaviour, changed user-facing capability, not representable as a tag, and three to five deployed models). This week supplied the vendor count; the architectural-distinction test is the open one, because a classifier head on a shared base may be a fine-tuning recipe rather than a new build.
Registry movement
Gemini 4 Argon carries status gated and placement confidence medium because the public cannot run it. MiMo-V2.6-Pro lifts last week's deferral: the Hugging Face repo now carries MIT license metadata and a model card. Grok 4.7's row now carries a 500K context, which the Amazon Bedrock listing prints citing xAI's September 21 launch announcement; last week's row had no context figure because the house had not recorded one.
- Added
- claude-sonnet-5-5, gpt-6-1-sol, gemini-4-argon, mimo-v2-6-pro, holo-4, kolibri-1
- Updated
- grok-4-7
Frontier movements
Anthropic · 2026-09-28 · Frontier · Reasoning
Claude Sonnet 5.5
Artificial Analysis places Sonnet 5.5 (max) at 56, two points behind Opus 5.5 at $2 / $10 against $4 / $20. It leads Terminal-Bench 4.0 on both the vendor harness and Artificial Analysis's independent run, 6.6 points apart (the table is in the benchmark moves). The visible fallback to Sonnet 5 on higher-risk cyber requests, driven by classifiers in the request path, is a behaviour your security team should test, because it changes which model answers mid-session; it is also the week's first-party example of a classifier sitting in a production approval path, which Agent Techniques picks up.
- Model registry ID
- claude-sonnet-5-5
Sources Anthropic
OpenAI · 2026-09-29 · Frontier · Reasoning
GPT-6.1 Sol
Sol has the cheapest cache meter at the tier and the only long-context surcharge (the meter table is in the architecture watch). OpenAI's DeepSWE v1.1 75.2% and $5.47 per Terminal-Bench Science task are vendor-run, and the cost-per-task advantage over Astra in OpenAI's table follows from Sol's list price being one-fifth of Astra's. Artificial Analysis's independent 52 puts Sol one point below Astra. Sol is classified Critical in cybersecurity (OpenAI's highest internal capability tier, which triggers its strictest deployment safeguards) and ships under Astra's safeguards stack, so expect the same refusal surface as the top tier. It replaced GPT-6 Sol after seven days; pin model ids.
- Model registry ID
- gpt-6-1-sol
Sources OpenAI
Google DeepMind · 2026-09-30 · Frontier · Reasoning
Gemini 4 Argon
Artificial Analysis scored Argon (high) at 53, equal to Astra, at about 60% of Astra's cost per task, and attributes the gain to two things: lower hallucinations and stronger agentic results (AutomationBench-AA 78%, first). The hallucination half rests partly on abstention. Per AA's September 30 article, Argon's 15% rate on AA-Omniscience (AA's test of whether a model answers wrongly or declines when it does not know) is the lowest of the three models AA compared (GPT-6 Astra 51%, GPT-6.1 Sol 54%; Sonnet 5.5 was not in that comparison and Astra is not a week's release), but its accuracy is 50% against Astra's 63%, so it declines more rather than knowing more. Google's table leads DeepSWE v1.1 at 77.9% (vendor-run) and trails Opus 5.5 on Terminal-Bench 4.0. The $2 / $10 rate is a 50% discount for at least one month, after which the card is $4 / $20. Access is limited to Fairwind Program cyber defenders; developers are 'next' with no date; note the 1M-token output limit for long-document generation when it arrives.
- Model registry ID
- gemini-4-argon
Sources Google
SpaceXAI · 2026-09-28 · Frontier · Reasoning
Grok 4.7
For AWS-committed enterprises this is a way to run a SpaceXAI model under existing Bedrock IAM, logging and private-link controls. AWS's listing prints the 500K context figure citing xAI's September 21 launch announcement, which the house had not recorded; the price on Bedrock should be checked against the $2 / $6 API rate before assuming parity.
- Model registry ID
- grok-4-7
Sources AWS Machine Learning Blog
Open weights
Xiaomi · 2026-09-21 · Open Frontier · Moe
MiMo-V2.6-Pro
Last week's deferral lifts: the repo carries MIT metadata and a model card for a 1.02T / 42B-active multimodal MoE ('active' is the parameters used per token). Artificial Analysis lists it at 46, ahead of GLM-5.3 (45) and Kimi K3 (44) and tied with Grok 4.7; its cost and minutes per task are in the benchmark moves. For a buyer the number to weigh is the latency: the cheapest open frontier model is also one of the slowest, which decides whether it fits interactive or batch work.
- Model registry ID
- mimo-v2-6-pro
Sources Hugging Face model card
Aleph Alpha · 2026-10-03 · Open Frontier · Moe
Kolibri-1
The fully open training recipe (20T pretraining tokens in 21 days, about 392k GPU-hours, roughly 6.4e23 FLOPs) is the useful artifact for anyone budgeting a sovereign model; the benchmark table (AIME 2025 96.9%, SWE-Bench Verified 66.4%) is vendor-run with no independent evaluation. At about 78 GB in FP8 (8-bit floating-point weights, half the memory of 16-bit) it fits a single H200 or B200; among European open-weight releases it is the first the house has recorded that combines a single-GPU footprint with a 1M context, though Mistral has shipped Apache 2.0 models with on-prem footprints at shorter contexts.
- Model registry ID
- kolibri-1
Sources Aleph Alpha
H Company · 2026-09-28 · Specialist · Agentic
Holo4
H Company's own figures are 45.4% on AutomationBench (business-workflow automation in simulated apps) at $0.05 per task for the 27B, read from its chart, and 61.7% on OSWorld 2.0 (desktop control) at $1.22 per task against Opus 5.5's 81.8%. The per-task costs are cheap in absolute terms; no frontier cost per AutomationBench or OSWorld task exists in the sources, so there is no frontier cost comparison to draw, and the vendor tables are from different versions. CC BY-NC 4.0 is more restrictive than the Apache 2.0 Qwen base, so enterprises need a commercial license from H Company before production use. Publishing all trajectories is the right precedent and lets you audit the score.
- Model registry ID
- holo-4
Sources Hugging Face blog (H Company)
Cloudflare · 2026-10-01 · Specialist · Dense
Clef
Clef returns calibrated probabilities over 1-64 typed questions in one forward pass, which is the shape an agent's allow/flag/block gate needs. Cloudflare's benchmark table (BANKING77 94.20 macro-F1, the average of per-class accuracy scores, against Jev's 79.74) is vendor-run, and the one community head-to-head posted on Hacker News, a 250-sample test, found Clef 5.2x more expensive and roughly five times slower than Jev for a 1.2-point gain. Read it as a reviewer model you can self-host under Apache 2.0, not as a leaderboard result; the technique read and the decision-model price list are in Agent Techniques.
Sources Cloudflare Blog
Architecture watch
List-price convergence with meter divergence
Three labs chose the same headline number and set every other meter independently. This is the fact home for the mid-tier meter packet. Cache reads (re-sending context the provider already holds): GPT-6.1 Sol $0.10 per million, Sonnet 5.5 $0.20, Argon 95% off list. Long-context repricing: Sol moves to $4 / $15 above 272K tokens; Sonnet 5.5 has no surcharge across its 1M context. Introductory pricing: Argon's $2 / $10 is a 50% discount for at least one month, reverting to $4 / $20, and nobody outside the Fairwind Program can pay either rate yet. As a heuristic rather than a measurement (the house has no token-mix data for production agents), the cache-read rate is the meter most likely to dominate agent sessions, which tend to reread context more than they emit tokens; the long-context threshold catches document and codebase workloads, and only Sol has it; the introductory reversion catches budget planning, and only Argon has it. A router that keys on list price will pick the wrong model for most agent workloads; key it on cache-read rate, context threshold and effective price after the intro period ends, and measure your own token mix to replace the heuristic.
- Examples
- Claude Sonnet 5.5 ($2 / $10, cache $0.20, no long-context tier), GPT-6.1 Sol ($2 / $10, cache $0.10, $4 / $15 above 272K), Gemini 4 Argon ($2 / $10 intro, $4 / $20 list, cached 95% off)
Sources OpenAI API pricing; Anthropic pricing; Google Gemini 4 Argon announcement
Task heads on a shared 27B base
Four releases in one week post-trained the same Qwen3.8-27B checkpoint into a specialist: three decision heads and one computer-use agent. The pattern is that a 27B dense open base has become the default substrate for task-specific heads, in the way BERT-base once was for classifiers, and that the differentiation is in the head, the loss and the data rather than the trunk. For buyers this means provenance and training-data disclosure matter more than parameter counts, and the Perplexity case, where the published shards match a community checkpoint byte for byte, shows how thin the provenance can be. For the tree it raises the question of whether decision models earn a branch, which the placement rules answer with a 3-of-6 independent-vendor test that this week put in play.
- Examples
- Cloudflare Clef (Qwen3.8-27B), Perplexity pplx-decider-v1-27b (Qwen3.8-27B), AutoTrust JEV-27B (Qwen3.8-27B), Holo4-27B (Qwen3.8-27B)
Sources Cloudflare, Perplexity, AutoTrust and H Company model cards on Hugging Face
Gating as a release stage
Each of the three labs shipped or announced a model this week with a safety gate that changes what a buyer actually gets. Google's is an access gate: the model exists, scores on independent indices, and is unavailable. Anthropic's is a routing gate: the model you called may hand the request to its predecessor mid-session, visibly, on the decision of reasoning-extraction classifiers in the request path. OpenAI's is a classification gate: Sol inherits Astra's Critical-cyber safeguards, and the planned Astra successor was withheld entirely. The architectural consequence is that 'model' now denotes a policy-wrapped system whose behaviour depends on request content, and evaluation harnesses that assume a fixed model behind an endpoint will produce numbers that do not reproduce in production. Test with your real request mix, including the requests the gate is designed to catch. The same architecture, a classifier deciding what the acting model may do, is the week's technique in Agent Techniques.
- Examples
- Gemini 4 Argon (Fairwind Program only), Claude Sonnet 5.5 (visible fallback to Sonnet 5 on higher-risk cyber tasks), GPT-6.1 Sol (Critical in cyber, Astra's safeguards stack), GPT-6.1 Astra (withheld)
Sources Google Gemini 4 Argon announcement; Claude Sonnet 5.5 System Card; OpenAI Deployment Safety Hub
Verbosity and latency as hidden costs of open frontier scores
The top open-weights model on the independent index is also one of the slowest: Artificial Analysis clocks MiMo-V2.6-Pro at 19.5 minutes per Intelligence Index task at $0.13, roughly a fifteenth of GLM-5.3's $2.01 per task for one more index point. AA's cost per task is price times tokens emitted, so a model that scores by generating long reasoning traces can be cheap per token and slow per task at the same time, and the comparison that matters is cost and minutes per task, not list price. Kolibri-1's 3.46B active on 78B total is the extreme case of a small active footprint, with no independent latency or tokens-per-task figure yet. An architect choosing an open model for interactive work should weight tokens per second and tokens per task as heavily as the index score; for batch work the index-per-dollar figure is the right one.
- Examples
- MiMo-V2.6-Pro (46 on the index, 19.5 minutes per task, $0.13 per task), GLM-5.3 (45 on the index, $2.01 per task), Kolibri-1 (3.46B active, vendor-run only)
Sources Artificial Analysis model pages
Benchmark moves
Artificial Analysis Intelligence Index v4.3.2
Four new entries in one week; Opus 5.5 holds first, Sonnet 5.5 enters second, Argon ties Astra, MiMo-V2.6-Pro confirmed as top open weights. The setting in parentheses is the reasoning-effort level the model was run at; scores are not comparable across settings
- Claude Opus 5.5 (max)
- 58
- Claude Sonnet 5.5 (max)
- 56
- Gemini 4 Argon (high)
- 53
- GPT-6 Astra (max)
- 53
- GPT-6.1 Sol (max)
- 52
Sources Artificial Analysis
Artificial Analysis Intelligence Index v4.3.2, open weights
MiMo-V2.6-Pro confirmed first among open-weights models, tying Grok 4.7 at a fraction of the cost per task; this table is the fact home for the open-weights cost and latency figures
- MiMo-V2.6-Pro
- 46 ($0.13 per task, 19.5 minutes per task, both read from AA's chart)
- GLM-5.3 (max)
- 45 ($2.01 per task, read from AA's chart)
- Kimi K3 (max)
- 44
- Grok 4.7 (closed, for reference)
- 46 ($2.73 per task, read from AA's chart)
Sources Artificial Analysis
Terminal-Bench 4.0 (independent Artificial Analysis runs, with the vendor figure for comparison)
Artificial Analysis's independent run puts Sonnet 5.5 first at 64%, 6.6 points below Anthropic's own 70.6%; Opus 5.5, Astra and Argon follow within seven points. The AA figures are read from AA's September 30 Argon article and sit outside the week's graded research set; the vendor figures are graded
- Claude Sonnet 5.5 (max) (Artificial Analysis, independent)
- 64%
- Claude Opus 5.5 (max) (Artificial Analysis, independent)
- 60%
- GPT-6 Astra (Artificial Analysis, independent)
- 59%
- Gemini 4 Argon (Artificial Analysis, independent; Google's own run 57.4%)
- 57%
- Claude Sonnet 5.5 (Anthropic, vendor-run, for comparison)
- 70.6%
DeepSWE v1.1 (vendor-run tables)
Argon claims the top score; Sol matches Astra at one-fifth the price; every figure is from the vendor's own table
- Gemini 4 Argon (Google)
- 77.9%
- GPT-6.1 Sol at high (OpenAI)
- 75.2%
- GPT-6 Astra (OpenAI / Google tables)
- ~74.8% / 74.1%
- Claude Opus 5.5 (Google table)
- 74.2%
- GPT-6 Sol best (OpenAI)
- 68.8%
Sources OpenAI; Google DeepMind
AutomationBench (vendor-run, mixed versions)
Argon leads on Google's table, 8.8 points ahead of Opus 5.5 on the same table; Holo4-27B reaches 45.4% at $0.05 per task on H Company's chart; Sol's 31.7% on OpenAI's table is a different version and is not comparable with Google's rows
- Gemini 4 Argon (Google table)
- 51.3%
- Holo4-27B (H Company, v1.0.6, read from chart)
- 45.4% at $0.05 per task
- Claude Opus 5.5 (Google table)
- 42.5%
- Holo4-35B-A3B (H Company, v1.0.6, read from chart)
- 34.5% at $0.02 per task
- GPT-6.1 Sol at medium (OpenAI table; OpenAI states a 2.2-point lead over Opus 5.5 on its own run, which is not the Google figure)
- 31.7%
LMArena text leaderboards (published Sep 30 and Oct 2)
Argon debuts first on Hard Prompts (preliminary) and Coding; MiMo-V2.6-Pro enters Coding at 1540
- Gemini 4 Argon, Coding
- 1560 ± 17
- Gemini 4 Argon, Hard Prompts (preliminary, 3,166 votes)
- 1551 ± 11
- MiMo-V2.6-Pro, Coding
- 1540 ± 18
Sources LMArena
Tier scorecard
As of 2026-10-03
| Tier | Leader | Challenger | Read |
|---|---|---|---|
| Closed frontier | Claude Opus 5.5 | GPT-6 Astra (tied on the index by Gemini 4 Argon, which cannot be bought) | Opus 5.5 leads the independent index at 58, five points clear of Astra and Argon at 53. Fable 5.1 held the slot through W39 on buyer-trace grounds (the house kept it as the default until a buyer-side trace showed Opus 5.5 matching it on real work); no such trace has arrived, and the house has no independent index score for Fable 5.1 in this window, so the handover rests on Opus 5.5's index lead and on Anthropic positioning Opus as its flagship, not on a measured Fable gap. Argon's tie with Astra is real and unpurchasable. |
| Open frontier | MiMo-V2.6-Pro | GLM-5.3 | Verified weights under MIT and an independent 46 make MiMo the open leader; GLM-5.3 at 45 is the challenger. DeepSeek-V4.1-Flash gives up the slot it held on serving cost; Kolibri-1 enters the watch column pending any independent score. |
| Reasoning | Claude Opus 5.5 | Claude Sonnet 5.5 | Two points separate them on the independent index at a 2x price difference. Sonnet 5.5 is the default for most reasoning workloads this quarter; Opus 5.5 is the ceiling. GPT-6.1 Sol at 52 is the cross-vendor alternative with the cheaper cache meter. |
| Coding | Claude Sonnet 5.5 | GPT-6.1 Sol | Sonnet 5.5 leads Terminal-Bench 4.0 on both the vendor and the independent run (table above) and takes the slot from Opus 5.5 at half the price; Sol replaces GPT-6 Astra as challenger because it matches Astra on OpenAI's DeepSWE table at one-fifth the cost. Argon's 77.9% DeepSWE is the highest claim and the least testable. |
| Multimodal | Gemini 3.8 Live Extended Thinking | Gemini 4 Argon | The buyable Google multimodal model holds the slot; Argon's 1M-token output and LVBench 91.7% (vendor-run) would take it when developers can call it. MiMo-V2.6-Pro's native text, image, video and audio input makes it the open alternative. |
| Edge / small | 4B decision heads via llama.cpp (Kev-4B and peers) | Clef-flash (9B) | This tier covers models meant to run on one device or one CPU, judged on a shipped task they do well rather than on the index. The day-0 decision heads, from 144 million parameters to 4B, answering allow/flag/block in 3-43 ms per question (author-reported; 3 ms is Julia-1 at 144 million parameters, and the 4B heads are 12-36 ms) through llama.cpp's typed endpoint take the slot from last week's leader, the vendor claim that a 30B mixture-of-experts model runs locally on the Snapdragon 8 Elite Extreme Gen 6 handset, which still has no independent throughput figure; Clef-flash at 9B under Apache 2.0 replaces Nex-N2.5 mini as challenger. Holo4-27B at $0.05 per AutomationBench task (vendor-run) is noted, but a 27B model served through an API is not an edge footprint. |
Vendor signals
2026-09-28 · OpenAI
Shelves GPT-6.1 Astra after internal alignment tests; the pause on an unreleased most-capable tier remains; publishes a 'towards safety cases' post with no numbers
GPT-6 Astra remains in production behind dots and the Ultrafast tier; what OpenAI withheld is its successor, and what it paused is a tier above Astra that never shipped. Architects should keep the cross-vendor fallback for top-tier tool use through Q4 because the roadmap above Astra is now empty. The post's absence of quantified results means the external evidence for the decision is the UK AI Security Institute's evaluation of GPT-6 Astra, which is about the model that stayed; the three-model table and its caveats are in Agent Techniques.
Sources OpenAI
2026-09-29 · OpenAI
Adds an Ultrafast service tier for GPT-6 Astra at $60 / $300 per million (6x) for up to 8x token speed in Codex and 6x in the API, per OpenAI; introduces Pro 500 at $500 a month; halves the Pro 200 Codex allowance from October 30
This is the first retail price increase on a frontier coding tier since the mid-tier convergence, and it prices speed as a premium product rather than a model property. Teams that depend on Codex throughput should model the move to Pro 500 or Sol Ultrafast against moving the workload to Sonnet 5.5 before October 30.
Sources OpenAI API pricing
2026-09-30 · Google DeepMind
Announces Gemini 4 Argon with access limited to Fairwind Program cyber defenders; publishes introductory pricing with no end date and no developer availability date
Google chose to claim the frontier on independent indices before letting anyone buy the model. For a buyer the signal is that Google's Q4 roadmap is public and its Q4 product is not. The reversion to $4 / $20 after the introductory month means the $2 / $10 figure in comparison pieces is a promotional rate, not Google's price.
Sources Google
2026-09-28 · Anthropic
Ships Sonnet 5.5 to all platforms including free Claude.ai; sets retirement no sooner than September 28, 2027; Haiku 5.5 slips to 'coming weeks' for a second week
A one-year retirement floor is a deployment guarantee most labs do not give and should go into vendor-selection scorecards. The Haiku slip matters for the cheap tier: there is still no Anthropic model below $2 / $10 in the 5.5 family, which leaves the sub-dollar tier to GPT-6 Luna, Gemini Flash and the open models.
Sources Claude Platform docs
2026-10-01 · Perplexity
Launches a Decisions API at $0.04 per million input tokens with free output, backed by pplx-decider-v1-27b; community analysis finds the weights byte-identical to a 12-day-old community checkpoint that the launch post does not mention
The price is the lowest in the decision-model category and the provenance is the weakest. Enterprises evaluating per-action review should treat the vendor benchmark table as describing someone else's model until Perplexity explains the match in its launch materials, and should prefer a vendor whose training provenance is documented.
Sources Hugging Face (perplexity-ai/pplx-decider-v1-27b); ModelSystem.One
2026-10-01 · Black Forest Labs
Launches FLUX 3 Image via API and a commercial-weights license, with open weights promised 'in the coming weeks'
The open-weights promise is a calendar item, not a release; FLUX 3 is tracked here only because its weights, when published, would be the first open image model of the generation. Until then it is an API product.
Sources Black Forest Labs
Watchlist
October
Claude Haiku 5.5 model id on the Claude Platform models page
Two weeks of 'coming weeks'. Haiku 5.5 would be Anthropic's first sub-$2 model with the 5.5 safeguards and the obvious reviewer model for per-action gating.
Q4 2026
Gemini 4 Argon developer availability and the introductory pricing end date
Independent benchmarks on a model nobody can run are a preview, not a result. Developer access resets the top-tier comparison and the $4 / $20 reversion tests the mid-tier convergence.
October
A second independent Terminal-Bench 4.0 harness on Claude Sonnet 5.5
Anthropic's 70.6% and Artificial Analysis's 64% are 6.6 points apart on the number most likely to move coding-agent defaults this quarter. A second independent harness would show whether the gap is AA's setup or the vendor's.
October
Perplexity response on pplx-decider-v1-27b provenance
Byte-identical shards to a community checkpoint, unmentioned in the launch post, is either an oversight or a disclosure problem. Which one it is decides whether the Decisions API belongs in an enterprise evaluation.
By Oct 30
Nvidia InferenceX submission for Vera Rubin NVL72
The lapsed Q3 commitment and its consequences are tracked in AI Stack Weekly; a Rubin submission would give the first independent cost-per-million figures for serving this week's models.
Q4 2026
First independent evaluation of Kolibri-1 and FLUX 3 open weights
Kolibri-1 pairs a single-GPU footprint with a 1M context, a combination no other European open-weight release in the house record has; its table is vendor-run. FLUX 3's weights decide whether the open image tier moves this generation.
Changelog
- Authored for the September 28 to October 4, 2026 window from vendor launch posts, model cards and Artificial Analysis and LMArena publications.
- Six tree rows added and one updated; the decision-model fine-tunes are reviewed and deferred pending the placement rules' 3-of-6 test rather than placed.
- Scorecard: Opus 5.5 takes the closed-frontier and reasoning leader slots from Fable 5.1 on the independent index; MiMo-V2.6-Pro takes the open leader slot from DeepSeek-V4.1-Flash on verified weights and an independent score; the Open frontier challenger moves from Atria Dawn Preview to GLM-5.3; the Coding leader and challenger move from Opus 5.5 and GPT-6 Astra to Sonnet 5.5 and GPT-6.1 Sol; the Edge / small leader moves from the Snapdragon 8 Elite Extreme on-device claim to the 4B llama.cpp decision heads and the challenger from Nex-N2.5 mini to Clef-flash; the Multimodal challenger moves from Gemini 3.8 Flash TTS to Gemini 4 Argon.
- Every benchmark figure is labelled vendor-run or independent at the point of use; vendor tables from different labs are not compared against each other except where flagged as such.
- Revision cycle 1 (editorial board): the Terminal-Bench 4.0 table now carries Artificial Analysis's independent runs for Sonnet 5.5 (64%), Opus 5.5 (60%), GPT-6 Astra (59%) and Argon (57%), which the first draft missed; Argon's hallucination figure is corrected from highest to lowest in the set, achieved by abstention; Argon's tier is corrected to frontier; the claim that OpenAI attributes Sol's cost advantage to cache reads is removed; the Opus 5.5 AutomationBench row is labelled derived; the bigRead is restructured around the three meters.
- Revision cycle 2 (editorial board): the Opus 5.5 AutomationBench row now carries Google's printed 42.5% instead of a derived figure and the Sol row is marked non-comparable; Argon's Omniscience comparison is restricted to the models in Artificial Analysis's September 30 article; the AA Terminal-Bench and Omniscience figures are labelled as outside the week's graded research set at the point of use (the fact audit verified them against the article; research.json still needs rows for them); the NOTICE-file provenance clause is removed as unverified; the Grok 4.7 context note no longer says 'undisclosed'; the Holo4 cost claim is narrowed to absolute per-task cost with no frontier comparison; the bigRead leads with the gating thesis, credits DataCamp, TokenCost and synthorai for the meter frame, and drops the 'house contribution' line; the mid-tier meter packet's fact home is the first architecture-watch entry and the AISI table's fact home is Agent Techniques; the Edge / small tier is defined and re-ranked; the Coding and Open frontier changes are recorded in the scorecard changelog line.