{
  "_meta": {
    "publication": "The Model Pulse",
    "schemaVersion": "2026.05.02",
    "generatedAt": "2026-10-04T02:02:09.016Z",
    "canonicalUrl": "https://brianletort.ai/industry/models/2026-W40",
    "markdownUrl": "https://brianletort.ai/industry/models/2026-W40/llm.md",
    "pdfUrl": "https://brianletort.ai/downloads/model-pulse-2026-W40.pdf",
    "treeUrl": "https://brianletort.ai/industry/tree",
    "sourceFile": "src/data/industry/models/2026-W40.ts"
  },
  "issue": {
    "slug": "2026-W40",
    "isoYear": 2026,
    "isoWeek": 40,
    "issueNumber": 24,
    "publishedAt": "2026-10-03",
    "cadence": "weekly",
    "periodLabel": "Week 40 of 2026",
    "bigRead": {
      "headline": "Same list price, three different bills",
      "body": "A 'model' is now a policy-wrapped system, and its price is a set of meters rather than a list rate: that is what this week's three mid-tier releases show, and it changes how you evaluate and how you budget. Start with the gates, because each lab shipped a different kind. Google's is an access gate: Gemini 4 Argon exists, scores on independent indices, and cannot be bought. Anthropic's is a routing gate: Sonnet 5.5 can hand a higher-risk cyber request to Sonnet 5 mid-session, visibly, through classifiers in the request path. OpenAI's is a classification gate: Sol inherits GPT-6 Astra's safeguards stack with a Critical rating in cybersecurity (OpenAI's highest internal capability tier, which triggers its strictest deployment controls), and the planned Astra successor was withheld with no numbers. The September 25 pause covers an unreleased most-capable tier; GPT-6 Astra itself still serves dots and the Ultrafast tier. An evaluation harness (the software that wraps the model, supplies its tools and runs the benchmark) that assumes a fixed model behind an endpoint will produce numbers that do not reproduce in production.\n\nThe meters are the second half. For a session that rereads a large context many times, the cache-read meter alone can make the Sol and Sonnet 5.5 bills diverge by 2x behind the same $2 / $10 list price; for a single long-document pass, Sol's long-context repricing does the opposite. Argon's $2 / $10 is a 50% introductory discount for at least one month, and TokenCost noted that its post-intro card is Opus 5.5's card line for line; DataCamp and TokenCost both published the meters-over-list-price read before this issue, and synthorai added the multiple the meter table does not show, that Sonnet 5.5 writes 1.7-3.1x the output tokens of Sol on comparable tasks. The full per-meter table, with the cache-read rates and the 272K threshold, is in the architecture watch below. Above the mid tier on the independent index, Opus 5.5 at $4 / $20 is the only buyable model that outscores Sonnet 5.5, and the question for most agent budgets is whether its two extra index points are worth doubling the bill; GPT-6 Astra at $10 / $50 is buyable and scores below Sonnet 5.5, which is the stronger point for a buyer.\n\nOne number to carry: Anthropic's own Terminal-Bench 4.0 harness puts Sonnet 5.5 at 70.6% and Artificial Analysis's independent run (its September 30 Argon article) at 64%, a 6.6-point gap that is the size of the vendor-versus-independent difference to expect in your own evaluation; the independent table, with the effort settings each model was run at, is in the benchmark moves.\n\nWhat to do: rerun the agent cost model with cache-read rates, long-context thresholds and output tokens per task as inputs, not list price; benchmark Sonnet 5.5 against Sol on your own task set rather than theirs, with the real request mix including the requests the gates are designed to catch; and treat Argon as a Q4 evaluation item, not a Q4 deployment option. The capital and power consequences are in AI Stack Weekly; the permission-layer consequences for agents are in Agent Techniques."
    },
    "treeDelta": {
      "summary": "Six rows added, one updated. Three closed frontier-adjacent models at $2 / $10 (Claude Sonnet 5.5, GPT-6.1 Sol, Gemini 4 Argon as gated), two open-weight MoE releases (mixture-of-experts, a model that activates only a few of its sub-networks per token; MiMo-V2.6-Pro now verified, Kolibri-1), and Holo4 under the Holo3.1 precedent (W23 placed H Company's computer-use family as a tree node because it ships in multiple sizes with quantized checkpoints, reduced-precision weights, for local inference). The decision-model fine-tunes are reviewed and deferred rather than placed: the tree's 3-of-6 rule creates a new branch only when at least three of six tests hold (durable architectural distinction, three or more independent vendors, changed infrastructure behaviour, changed user-facing capability, not representable as a tag, and three to five deployed models). This week supplied the vendor count; the architectural-distinction test is the open one, because a classifier head on a shared base may be a fine-tuning recipe rather than a new build.",
      "added": [
        "claude-sonnet-5-5",
        "gpt-6-1-sol",
        "gemini-4-argon",
        "mimo-v2-6-pro",
        "holo-4",
        "kolibri-1"
      ],
      "updated": [
        "grok-4-7"
      ],
      "reviewed": [
        {
          "name": "GPT-6.1 Astra",
          "disposition": "excluded",
          "reason": "Withheld by OpenAI before release on alignment grounds; the tree records shipped models, and no model id, card or price exists for it."
        },
        {
          "name": "GPT-6 Sol",
          "disposition": "excluded",
          "reason": "Replaced by GPT-6.1 Sol seven days after it shipped. Its existence and 68.8% DeepSWE best are recorded in the gpt-6-1-sol row rather than as a separate node."
        },
        {
          "name": "GPT-6 Astra Ultrafast",
          "disposition": "excluded",
          "reason": "A service tier for an existing model at 6x the price for up to 8x the speed, not a new model; carried as a vendor signal in this issue."
        },
        {
          "name": "Clef",
          "disposition": "deferred",
          "reason": "Cloudflare's 27B and 9B decision models are typed-classifier heads (a small classifier layer trained on top of a general model, returning probabilities over fixed options rather than text) post-trained on Qwen3.8-27B and Qwen3.5-9B. Deferred until the placement rules' 3-of-6 test is applied to decision models as a category; three independent vendors shipped one this week, so the test is live."
        },
        {
          "name": "pplx-decider-v1-27b",
          "disposition": "excluded",
          "reason": "ModelSystem.One reports the published shards are byte-identical to the community denis-pplx/autojev-27b checkpoint and the launch post does not mention it. Provenance must be resolved before the tree carries a row."
        },
        {
          "name": "JEV-27B",
          "disposition": "deferred",
          "reason": "AutoTrust's Apache 2.0 LoRA decision head (a small set of adapter weights plus a classifier layer) on Qwen3.8-27B, distilled from Jev 1.13 outputs. Same category question as Clef; deferred with it."
        },
        {
          "name": "llama.cpp day-zero decision-model GGUFs (Julia-1, Laya, Kev-4B, lev, OpenJev 27B)",
          "disposition": "excluded",
          "reason": "Small task heads converted to GGUF (llama.cpp's single-file weight format for local inference) for the new /v1/systemone endpoint, without model cards of their own; the endpoint is the news and it is covered in Agent Techniques."
        },
        {
          "name": "FLUX 3 Image",
          "disposition": "excluded",
          "reason": "Image-generation model served via API with weights promised 'in the coming weeks'; the tree does not carry image-only models and the weights were not public in the window."
        },
        {
          "name": "Ideogram 4.5",
          "disposition": "excluded",
          "reason": "Image-generation model update with no open weights; outside the tree's scope."
        },
        {
          "name": "MiMo-V2.6-Flash",
          "disposition": "deferred",
          "reason": "Shares the Pro repository with a 309B / 15B-active configuration and a score on the Vals Index (a third-party leaderboard), but has no separate card or independent Intelligence Index placement yet; noted in the Pro row."
        }
      ],
      "note": "Gemini 4 Argon carries status gated and placement confidence medium because the public cannot run it. MiMo-V2.6-Pro lifts last week's deferral: the Hugging Face repo now carries MIT license metadata and a model card. Grok 4.7's row now carries a 500K context, which the Amazon Bedrock listing prints citing xAI's September 21 launch announcement; last week's row had no context figure because the house had not recorded one."
    },
    "frontierMovements": [
      {
        "modelId": "claude-sonnet-5-5",
        "name": "Claude Sonnet 5.5",
        "vendor": "Anthropic",
        "releaseDate": "2026-09-28",
        "headline": "Second on the independent index at half the price of first, and the first Sonnet with frontier cyber safeguards",
        "why": "Artificial Analysis places Sonnet 5.5 (max) at 56, two points behind Opus 5.5 at $2 / $10 against $4 / $20. It leads Terminal-Bench 4.0 on both the vendor harness and Artificial Analysis's independent run, 6.6 points apart (the table is in the benchmark moves). The visible fallback to Sonnet 5 on higher-risk cyber requests, driven by classifiers in the request path, is a behaviour your security team should test, because it changes which model answers mid-session; it is also the week's first-party example of a classifier sitting in a production approval path, which Agent Techniques picks up.",
        "tier": "frontier",
        "architecture": "reasoning",
        "source": "Anthropic",
        "sourceUrl": "https://www.anthropic.com/claude-sonnet-5-5"
      },
      {
        "modelId": "gpt-6-1-sol",
        "name": "GPT-6.1 Sol",
        "vendor": "OpenAI",
        "releaseDate": "2026-09-29",
        "headline": "Near-Astra on OpenAI's tables at one-fifth the price, with the cheapest cache meter at the tier",
        "why": "Sol has the cheapest cache meter at the tier and the only long-context surcharge (the meter table is in the architecture watch). OpenAI's DeepSWE v1.1 75.2% and $5.47 per Terminal-Bench Science task are vendor-run, and the cost-per-task advantage over Astra in OpenAI's table follows from Sol's list price being one-fifth of Astra's. Artificial Analysis's independent 52 puts Sol one point below Astra. Sol is classified Critical in cybersecurity (OpenAI's highest internal capability tier, which triggers its strictest deployment safeguards) and ships under Astra's safeguards stack, so expect the same refusal surface as the top tier. It replaced GPT-6 Sol after seven days; pin model ids.",
        "tier": "frontier",
        "architecture": "reasoning",
        "source": "OpenAI",
        "sourceUrl": "https://openai.com/index/introducing-gpt-6-1-sol/"
      },
      {
        "modelId": "gemini-4-argon",
        "name": "Gemini 4 Argon",
        "vendor": "Google DeepMind",
        "releaseDate": "2026-09-30",
        "headline": "Google's first above-Flash model in over seven months (Artificial Analysis's line) ties GPT-6 Astra on the independent index and cannot be bought",
        "why": "Artificial Analysis scored Argon (high) at 53, equal to Astra, at about 60% of Astra's cost per task, and attributes the gain to two things: lower hallucinations and stronger agentic results (AutomationBench-AA 78%, first). The hallucination half rests partly on abstention. Per AA's September 30 article, Argon's 15% rate on AA-Omniscience (AA's test of whether a model answers wrongly or declines when it does not know) is the lowest of the three models AA compared (GPT-6 Astra 51%, GPT-6.1 Sol 54%; Sonnet 5.5 was not in that comparison and Astra is not a week's release), but its accuracy is 50% against Astra's 63%, so it declines more rather than knowing more. Google's table leads DeepSWE v1.1 at 77.9% (vendor-run) and trails Opus 5.5 on Terminal-Bench 4.0. The $2 / $10 rate is a 50% discount for at least one month, after which the card is $4 / $20. Access is limited to Fairwind Program cyber defenders; developers are 'next' with no date; note the 1M-token output limit for long-document generation when it arrives.",
        "tier": "frontier",
        "architecture": "reasoning",
        "source": "Google",
        "sourceUrl": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"
      },
      {
        "modelId": "grok-4-7",
        "name": "Grok 4.7",
        "vendor": "SpaceXAI",
        "releaseDate": "2026-09-28",
        "headline": "Lands on Amazon Bedrock with a 500K context and four reasoning-effort levels, the first hyperscaler listing to print the context figure",
        "why": "For AWS-committed enterprises this is a way to run a SpaceXAI model under existing Bedrock IAM, logging and private-link controls. AWS's listing prints the 500K context figure citing xAI's September 21 launch announcement, which the house had not recorded; the price on Bedrock should be checked against the $2 / $6 API rate before assuming parity.",
        "tier": "frontier",
        "architecture": "reasoning",
        "source": "AWS Machine Learning Blog",
        "sourceUrl": "https://aws.amazon.com/blogs/machine-learning/grok-4-7-is-now-available-on-amazon-bedrock/"
      }
    ],
    "openWeights": [
      {
        "modelId": "mimo-v2-6-pro",
        "name": "MiMo-V2.6-Pro",
        "vendor": "Xiaomi",
        "releaseDate": "2026-09-21",
        "headline": "Weights verified under MIT; independent trackers confirm it as the top open-weights model",
        "why": "Last week's deferral lifts: the repo carries MIT metadata and a model card for a 1.02T / 42B-active multimodal MoE ('active' is the parameters used per token). Artificial Analysis lists it at 46, ahead of GLM-5.3 (45) and Kimi K3 (44) and tied with Grok 4.7; its cost and minutes per task are in the benchmark moves. For a buyer the number to weigh is the latency: the cheapest open frontier model is also one of the slowest, which decides whether it fits interactive or batch work.",
        "tier": "open_frontier",
        "architecture": "moe",
        "source": "Hugging Face model card",
        "sourceUrl": "https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/README.md"
      },
      {
        "modelId": "kolibri-1",
        "name": "Kolibri-1",
        "vendor": "Aleph Alpha",
        "releaseDate": "2026-10-03",
        "headline": "A 78B / 3.46B-active English-German MoE under Apache 2.0 with a 1M context, trained on 768 B200s",
        "why": "The fully open training recipe (20T pretraining tokens in 21 days, about 392k GPU-hours, roughly 6.4e23 FLOPs) is the useful artifact for anyone budgeting a sovereign model; the benchmark table (AIME 2025 96.9%, SWE-Bench Verified 66.4%) is vendor-run with no independent evaluation. At about 78 GB in FP8 (8-bit floating-point weights, half the memory of 16-bit) it fits a single H200 or B200; among European open-weight releases it is the first the house has recorded that combines a single-GPU footprint with a 1M context, though Mistral has shipped Apache 2.0 models with on-prem footprints at shorter contexts.",
        "tier": "open_frontier",
        "architecture": "moe",
        "source": "Aleph Alpha",
        "sourceUrl": "https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/"
      },
      {
        "modelId": "holo-4",
        "name": "Holo4",
        "vendor": "H Company",
        "releaseDate": "2026-09-28",
        "headline": "Computer-use family at 27B dense and 35B-A3B with every benchmark trajectory published, under a non-commercial license",
        "why": "H Company's own figures are 45.4% on AutomationBench (business-workflow automation in simulated apps) at $0.05 per task for the 27B, read from its chart, and 61.7% on OSWorld 2.0 (desktop control) at $1.22 per task against Opus 5.5's 81.8%. The per-task costs are cheap in absolute terms; no frontier cost per AutomationBench or OSWorld task exists in the sources, so there is no frontier cost comparison to draw, and the vendor tables are from different versions. CC BY-NC 4.0 is more restrictive than the Apache 2.0 Qwen base, so enterprises need a commercial license from H Company before production use. Publishing all trajectories is the right precedent and lets you audit the score.",
        "tier": "specialist",
        "architecture": "agentic",
        "source": "Hugging Face blog (H Company)",
        "sourceUrl": "https://huggingface.co/blog/Hcompany/holo4"
      },
      {
        "name": "Clef",
        "vendor": "Cloudflare",
        "releaseDate": "2026-10-01",
        "headline": "Open-weight decision model on Workers AI at $0.24 per million input tokens, with a 9B Clef-flash sibling",
        "why": "Clef returns calibrated probabilities over 1-64 typed questions in one forward pass, which is the shape an agent's allow/flag/block gate needs. Cloudflare's benchmark table (BANKING77 94.20 macro-F1, the average of per-class accuracy scores, against Jev's 79.74) is vendor-run, and the one community head-to-head posted on Hacker News, a 250-sample test, found Clef 5.2x more expensive and roughly five times slower than Jev for a 1.2-point gain. Read it as a reviewer model you can self-host under Apache 2.0, not as a leaderboard result; the technique read and the decision-model price list are in Agent Techniques.",
        "tier": "specialist",
        "architecture": "dense",
        "source": "Cloudflare Blog",
        "sourceUrl": "https://blog.cloudflare.com/clef-decision-models/"
      }
    ],
    "architectureWatch": [
      {
        "pattern": "List-price convergence with meter divergence",
        "examples": [
          "Claude Sonnet 5.5 ($2 / $10, cache $0.20, no long-context tier)",
          "GPT-6.1 Sol ($2 / $10, cache $0.10, $4 / $15 above 272K)",
          "Gemini 4 Argon ($2 / $10 intro, $4 / $20 list, cached 95% off)"
        ],
        "body": "Three labs chose the same headline number and set every other meter independently. This is the fact home for the mid-tier meter packet. Cache reads (re-sending context the provider already holds): GPT-6.1 Sol $0.10 per million, Sonnet 5.5 $0.20, Argon 95% off list. Long-context repricing: Sol moves to $4 / $15 above 272K tokens; Sonnet 5.5 has no surcharge across its 1M context. Introductory pricing: Argon's $2 / $10 is a 50% discount for at least one month, reverting to $4 / $20, and nobody outside the Fairwind Program can pay either rate yet. As a heuristic rather than a measurement (the house has no token-mix data for production agents), the cache-read rate is the meter most likely to dominate agent sessions, which tend to reread context more than they emit tokens; the long-context threshold catches document and codebase workloads, and only Sol has it; the introductory reversion catches budget planning, and only Argon has it. A router that keys on list price will pick the wrong model for most agent workloads; key it on cache-read rate, context threshold and effective price after the intro period ends, and measure your own token mix to replace the heuristic.",
        "source": "OpenAI API pricing; Anthropic pricing; Google Gemini 4 Argon announcement",
        "sourceUrl": "https://developers.openai.com/api/docs/pricing"
      },
      {
        "pattern": "Task heads on a shared 27B base",
        "examples": [
          "Cloudflare Clef (Qwen3.8-27B)",
          "Perplexity pplx-decider-v1-27b (Qwen3.8-27B)",
          "AutoTrust JEV-27B (Qwen3.8-27B)",
          "Holo4-27B (Qwen3.8-27B)"
        ],
        "body": "Four releases in one week post-trained the same Qwen3.8-27B checkpoint into a specialist: three decision heads and one computer-use agent. The pattern is that a 27B dense open base has become the default substrate for task-specific heads, in the way BERT-base once was for classifiers, and that the differentiation is in the head, the loss and the data rather than the trunk. For buyers this means provenance and training-data disclosure matter more than parameter counts, and the Perplexity case, where the published shards match a community checkpoint byte for byte, shows how thin the provenance can be. For the tree it raises the question of whether decision models earn a branch, which the placement rules answer with a 3-of-6 independent-vendor test that this week put in play.",
        "source": "Cloudflare, Perplexity, AutoTrust and H Company model cards on Hugging Face",
        "sourceUrl": "https://huggingface.co/Cloudflare/clef"
      },
      {
        "pattern": "Gating as a release stage",
        "examples": [
          "Gemini 4 Argon (Fairwind Program only)",
          "Claude Sonnet 5.5 (visible fallback to Sonnet 5 on higher-risk cyber tasks)",
          "GPT-6.1 Sol (Critical in cyber, Astra's safeguards stack)",
          "GPT-6.1 Astra (withheld)"
        ],
        "body": "Each of the three labs shipped or announced a model this week with a safety gate that changes what a buyer actually gets. Google's is an access gate: the model exists, scores on independent indices, and is unavailable. Anthropic's is a routing gate: the model you called may hand the request to its predecessor mid-session, visibly, on the decision of reasoning-extraction classifiers in the request path. OpenAI's is a classification gate: Sol inherits Astra's Critical-cyber safeguards, and the planned Astra successor was withheld entirely. The architectural consequence is that 'model' now denotes a policy-wrapped system whose behaviour depends on request content, and evaluation harnesses that assume a fixed model behind an endpoint will produce numbers that do not reproduce in production. Test with your real request mix, including the requests the gate is designed to catch. The same architecture, a classifier deciding what the acting model may do, is the week's technique in Agent Techniques.",
        "source": "Google Gemini 4 Argon announcement; Claude Sonnet 5.5 System Card; OpenAI Deployment Safety Hub",
        "sourceUrl": "https://deploymentsafety.openai.com/gpt-6-1-sol/introduction"
      },
      {
        "pattern": "Verbosity and latency as hidden costs of open frontier scores",
        "examples": [
          "MiMo-V2.6-Pro (46 on the index, 19.5 minutes per task, $0.13 per task)",
          "GLM-5.3 (45 on the index, $2.01 per task)",
          "Kolibri-1 (3.46B active, vendor-run only)"
        ],
        "body": "The top open-weights model on the independent index is also one of the slowest: Artificial Analysis clocks MiMo-V2.6-Pro at 19.5 minutes per Intelligence Index task at $0.13, roughly a fifteenth of GLM-5.3's $2.01 per task for one more index point. AA's cost per task is price times tokens emitted, so a model that scores by generating long reasoning traces can be cheap per token and slow per task at the same time, and the comparison that matters is cost and minutes per task, not list price. Kolibri-1's 3.46B active on 78B total is the extreme case of a small active footprint, with no independent latency or tokens-per-task figure yet. An architect choosing an open model for interactive work should weight tokens per second and tokens per task as heavily as the index score; for batch work the index-per-dollar figure is the right one.",
        "source": "Artificial Analysis model pages",
        "sourceUrl": "https://artificialanalysis.ai/models/mimo-v2-6-pro"
      }
    ],
    "benchmarkMoves": [
      {
        "benchmark": "Artificial Analysis Intelligence Index v4.3.2",
        "movement": "Four new entries in one week; Opus 5.5 holds first, Sonnet 5.5 enters second, Argon ties Astra, MiMo-V2.6-Pro confirmed as top open weights. The setting in parentheses is the reasoning-effort level the model was run at; scores are not comparable across settings",
        "rows": [
          {
            "model": "Claude Opus 5.5 (max)",
            "score": "58"
          },
          {
            "model": "Claude Sonnet 5.5 (max)",
            "score": "56"
          },
          {
            "model": "Gemini 4 Argon (high)",
            "score": "53"
          },
          {
            "model": "GPT-6 Astra (max)",
            "score": "53"
          },
          {
            "model": "GPT-6.1 Sol (max)",
            "score": "52"
          }
        ],
        "source": "Artificial Analysis",
        "sourceUrl": "https://artificialanalysis.ai/changelog"
      },
      {
        "benchmark": "Artificial Analysis Intelligence Index v4.3.2, open weights",
        "movement": "MiMo-V2.6-Pro confirmed first among open-weights models, tying Grok 4.7 at a fraction of the cost per task; this table is the fact home for the open-weights cost and latency figures",
        "rows": [
          {
            "model": "MiMo-V2.6-Pro",
            "score": "46 ($0.13 per task, 19.5 minutes per task, both read from AA's chart)"
          },
          {
            "model": "GLM-5.3 (max)",
            "score": "45 ($2.01 per task, read from AA's chart)"
          },
          {
            "model": "Kimi K3 (max)",
            "score": "44"
          },
          {
            "model": "Grok 4.7 (closed, for reference)",
            "score": "46 ($2.73 per task, read from AA's chart)"
          }
        ],
        "source": "Artificial Analysis",
        "sourceUrl": "https://artificialanalysis.ai/models/mimo-v2-6-pro"
      },
      {
        "benchmark": "Terminal-Bench 4.0 (independent Artificial Analysis runs, with the vendor figure for comparison)",
        "movement": "Artificial Analysis's independent run puts Sonnet 5.5 first at 64%, 6.6 points below Anthropic's own 70.6%; Opus 5.5, Astra and Argon follow within seven points. The AA figures are read from AA's September 30 Argon article and sit outside the week's graded research set; the vendor figures are graded",
        "rows": [
          {
            "model": "Claude Sonnet 5.5 (max) (Artificial Analysis, independent)",
            "score": "64%"
          },
          {
            "model": "Claude Opus 5.5 (max) (Artificial Analysis, independent)",
            "score": "60%"
          },
          {
            "model": "GPT-6 Astra (Artificial Analysis, independent)",
            "score": "59%"
          },
          {
            "model": "Gemini 4 Argon (Artificial Analysis, independent; Google's own run 57.4%)",
            "score": "57%"
          },
          {
            "model": "Claude Sonnet 5.5 (Anthropic, vendor-run, for comparison)",
            "score": "70.6%"
          }
        ],
        "source": "Anthropic; Google DeepMind; Artificial Analysis",
        "sourceUrl": "https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs"
      },
      {
        "benchmark": "DeepSWE v1.1 (vendor-run tables)",
        "movement": "Argon claims the top score; Sol matches Astra at one-fifth the price; every figure is from the vendor's own table",
        "rows": [
          {
            "model": "Gemini 4 Argon (Google)",
            "score": "77.9%"
          },
          {
            "model": "GPT-6.1 Sol at high (OpenAI)",
            "score": "75.2%"
          },
          {
            "model": "GPT-6 Astra (OpenAI / Google tables)",
            "score": "~74.8% / 74.1%"
          },
          {
            "model": "Claude Opus 5.5 (Google table)",
            "score": "74.2%"
          },
          {
            "model": "GPT-6 Sol best (OpenAI)",
            "score": "68.8%"
          }
        ],
        "source": "OpenAI; Google DeepMind",
        "sourceUrl": "https://deepmind.google/models/gemini/"
      },
      {
        "benchmark": "AutomationBench (vendor-run, mixed versions)",
        "movement": "Argon leads on Google's table, 8.8 points ahead of Opus 5.5 on the same table; Holo4-27B reaches 45.4% at $0.05 per task on H Company's chart; Sol's 31.7% on OpenAI's table is a different version and is not comparable with Google's rows",
        "rows": [
          {
            "model": "Gemini 4 Argon (Google table)",
            "score": "51.3%"
          },
          {
            "model": "Holo4-27B (H Company, v1.0.6, read from chart)",
            "score": "45.4% at $0.05 per task"
          },
          {
            "model": "Claude Opus 5.5 (Google table)",
            "score": "42.5%"
          },
          {
            "model": "Holo4-35B-A3B (H Company, v1.0.6, read from chart)",
            "score": "34.5% at $0.02 per task"
          },
          {
            "model": "GPT-6.1 Sol at medium (OpenAI table; OpenAI states a 2.2-point lead over Opus 5.5 on its own run, which is not the Google figure)",
            "score": "31.7%"
          }
        ],
        "source": "Google DeepMind; H Company; OpenAI",
        "sourceUrl": "https://huggingface.co/blog/Hcompany/holo4"
      },
      {
        "benchmark": "LMArena text leaderboards (published Sep 30 and Oct 2)",
        "movement": "Argon debuts first on Hard Prompts (preliminary) and Coding; MiMo-V2.6-Pro enters Coding at 1540",
        "rows": [
          {
            "model": "Gemini 4 Argon, Coding",
            "score": "1560 ± 17"
          },
          {
            "model": "Gemini 4 Argon, Hard Prompts (preliminary, 3,166 votes)",
            "score": "1551 ± 11"
          },
          {
            "model": "MiMo-V2.6-Pro, Coding",
            "score": "1540 ± 18"
          }
        ],
        "source": "LMArena",
        "sourceUrl": "https://arena.ai/leaderboard/text/hard-prompts"
      }
    ],
    "scorecard": {
      "asOf": "2026-10-03",
      "rows": [
        {
          "tier": "Closed frontier",
          "leader": "Claude Opus 5.5",
          "challenger": "GPT-6 Astra (tied on the index by Gemini 4 Argon, which cannot be bought)",
          "note": "Opus 5.5 leads the independent index at 58, five points clear of Astra and Argon at 53. Fable 5.1 held the slot through W39 on buyer-trace grounds (the house kept it as the default until a buyer-side trace showed Opus 5.5 matching it on real work); no such trace has arrived, and the house has no independent index score for Fable 5.1 in this window, so the handover rests on Opus 5.5's index lead and on Anthropic positioning Opus as its flagship, not on a measured Fable gap. Argon's tie with Astra is real and unpurchasable."
        },
        {
          "tier": "Open frontier",
          "leader": "MiMo-V2.6-Pro",
          "challenger": "GLM-5.3",
          "note": "Verified weights under MIT and an independent 46 make MiMo the open leader; GLM-5.3 at 45 is the challenger. DeepSeek-V4.1-Flash gives up the slot it held on serving cost; Kolibri-1 enters the watch column pending any independent score."
        },
        {
          "tier": "Reasoning",
          "leader": "Claude Opus 5.5",
          "challenger": "Claude Sonnet 5.5",
          "note": "Two points separate them on the independent index at a 2x price difference. Sonnet 5.5 is the default for most reasoning workloads this quarter; Opus 5.5 is the ceiling. GPT-6.1 Sol at 52 is the cross-vendor alternative with the cheaper cache meter."
        },
        {
          "tier": "Coding",
          "leader": "Claude Sonnet 5.5",
          "challenger": "GPT-6.1 Sol",
          "note": "Sonnet 5.5 leads Terminal-Bench 4.0 on both the vendor and the independent run (table above) and takes the slot from Opus 5.5 at half the price; Sol replaces GPT-6 Astra as challenger because it matches Astra on OpenAI's DeepSWE table at one-fifth the cost. Argon's 77.9% DeepSWE is the highest claim and the least testable."
        },
        {
          "tier": "Multimodal",
          "leader": "Gemini 3.8 Live Extended Thinking",
          "challenger": "Gemini 4 Argon",
          "note": "The buyable Google multimodal model holds the slot; Argon's 1M-token output and LVBench 91.7% (vendor-run) would take it when developers can call it. MiMo-V2.6-Pro's native text, image, video and audio input makes it the open alternative."
        },
        {
          "tier": "Edge / small",
          "leader": "4B decision heads via llama.cpp (Kev-4B and peers)",
          "challenger": "Clef-flash (9B)",
          "note": "This tier covers models meant to run on one device or one CPU, judged on a shipped task they do well rather than on the index. The day-0 decision heads, from 144 million parameters to 4B, answering allow/flag/block in 3-43 ms per question (author-reported; 3 ms is Julia-1 at 144 million parameters, and the 4B heads are 12-36 ms) through llama.cpp's typed endpoint take the slot from last week's leader, the vendor claim that a 30B mixture-of-experts model runs locally on the Snapdragon 8 Elite Extreme Gen 6 handset, which still has no independent throughput figure; Clef-flash at 9B under Apache 2.0 replaces Nex-N2.5 mini as challenger. Holo4-27B at $0.05 per AutomationBench task (vendor-run) is noted, but a 27B model served through an API is not an edge footprint."
        }
      ]
    },
    "vendorSignals": [
      {
        "vendor": "OpenAI",
        "date": "2026-09-28",
        "signal": "Shelves GPT-6.1 Astra after internal alignment tests; the pause on an unreleased most-capable tier remains; publishes a 'towards safety cases' post with no numbers",
        "meaning": "GPT-6 Astra remains in production behind dots and the Ultrafast tier; what OpenAI withheld is its successor, and what it paused is a tier above Astra that never shipped. Architects should keep the cross-vendor fallback for top-tier tool use through Q4 because the roadmap above Astra is now empty. The post's absence of quantified results means the external evidence for the decision is the UK AI Security Institute's evaluation of GPT-6 Astra, which is about the model that stayed; the three-model table and its caveats are in Agent Techniques.",
        "source": "OpenAI",
        "sourceUrl": "https://openai.com/index/towards-safety-cases-for-frontier-ai-training/"
      },
      {
        "vendor": "OpenAI",
        "date": "2026-09-29",
        "signal": "Adds an Ultrafast service tier for GPT-6 Astra at $60 / $300 per million (6x) for up to 8x token speed in Codex and 6x in the API, per OpenAI; introduces Pro 500 at $500 a month; halves the Pro 200 Codex allowance from October 30",
        "meaning": "This is the first retail price increase on a frontier coding tier since the mid-tier convergence, and it prices speed as a premium product rather than a model property. Teams that depend on Codex throughput should model the move to Pro 500 or Sol Ultrafast against moving the workload to Sonnet 5.5 before October 30.",
        "source": "OpenAI API pricing",
        "sourceUrl": "https://developers.openai.com/api/docs/pricing"
      },
      {
        "vendor": "Google DeepMind",
        "date": "2026-09-30",
        "signal": "Announces Gemini 4 Argon with access limited to Fairwind Program cyber defenders; publishes introductory pricing with no end date and no developer availability date",
        "meaning": "Google chose to claim the frontier on independent indices before letting anyone buy the model. For a buyer the signal is that Google's Q4 roadmap is public and its Q4 product is not. The reversion to $4 / $20 after the introductory month means the $2 / $10 figure in comparison pieces is a promotional rate, not Google's price.",
        "source": "Google",
        "sourceUrl": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"
      },
      {
        "vendor": "Anthropic",
        "date": "2026-09-28",
        "signal": "Ships Sonnet 5.5 to all platforms including free Claude.ai; sets retirement no sooner than September 28, 2027; Haiku 5.5 slips to 'coming weeks' for a second week",
        "meaning": "A one-year retirement floor is a deployment guarantee most labs do not give and should go into vendor-selection scorecards. The Haiku slip matters for the cheap tier: there is still no Anthropic model below $2 / $10 in the 5.5 family, which leaves the sub-dollar tier to GPT-6 Luna, Gemini Flash and the open models.",
        "source": "Claude Platform docs",
        "sourceUrl": "https://platform.claude.com/docs/en/models/sonnet-5-5/overview"
      },
      {
        "vendor": "Perplexity",
        "date": "2026-10-01",
        "signal": "Launches a Decisions API at $0.04 per million input tokens with free output, backed by pplx-decider-v1-27b; community analysis finds the weights byte-identical to a 12-day-old community checkpoint that the launch post does not mention",
        "meaning": "The price is the lowest in the decision-model category and the provenance is the weakest. Enterprises evaluating per-action review should treat the vendor benchmark table as describing someone else's model until Perplexity explains the match in its launch materials, and should prefer a vendor whose training provenance is documented.",
        "source": "Hugging Face (perplexity-ai/pplx-decider-v1-27b); ModelSystem.One",
        "sourceUrl": "https://huggingface.co/perplexity-ai/pplx-decider-v1-27b"
      },
      {
        "vendor": "Black Forest Labs",
        "date": "2026-10-01",
        "signal": "Launches FLUX 3 Image via API and a commercial-weights license, with open weights promised 'in the coming weeks'",
        "meaning": "The open-weights promise is a calendar item, not a release; FLUX 3 is tracked here only because its weights, when published, would be the first open image model of the generation. Until then it is an API product.",
        "source": "Black Forest Labs",
        "sourceUrl": "https://bfl.ai/models/flux-3-image"
      }
    ],
    "watchlist": [
      {
        "window": "October",
        "title": "Claude Haiku 5.5 model id on the Claude Platform models page",
        "why": "Two weeks of 'coming weeks'. Haiku 5.5 would be Anthropic's first sub-$2 model with the 5.5 safeguards and the obvious reviewer model for per-action gating."
      },
      {
        "window": "Q4 2026",
        "title": "Gemini 4 Argon developer availability and the introductory pricing end date",
        "why": "Independent benchmarks on a model nobody can run are a preview, not a result. Developer access resets the top-tier comparison and the $4 / $20 reversion tests the mid-tier convergence."
      },
      {
        "window": "October",
        "title": "A second independent Terminal-Bench 4.0 harness on Claude Sonnet 5.5",
        "why": "Anthropic's 70.6% and Artificial Analysis's 64% are 6.6 points apart on the number most likely to move coding-agent defaults this quarter. A second independent harness would show whether the gap is AA's setup or the vendor's."
      },
      {
        "window": "October",
        "title": "Perplexity response on pplx-decider-v1-27b provenance",
        "why": "Byte-identical shards to a community checkpoint, unmentioned in the launch post, is either an oversight or a disclosure problem. Which one it is decides whether the Decisions API belongs in an enterprise evaluation."
      },
      {
        "window": "By Oct 30",
        "title": "Nvidia InferenceX submission for Vera Rubin NVL72",
        "why": "The lapsed Q3 commitment and its consequences are tracked in AI Stack Weekly; a Rubin submission would give the first independent cost-per-million figures for serving this week's models."
      },
      {
        "window": "Q4 2026",
        "title": "First independent evaluation of Kolibri-1 and FLUX 3 open weights",
        "why": "Kolibri-1 pairs a single-GPU footprint with a 1M context, a combination no other European open-weight release in the house record has; its table is vendor-run. FLUX 3's weights decide whether the open image tier moves this generation."
      }
    ],
    "changelog": [
      "Authored for the September 28 to October 4, 2026 window from vendor launch posts, model cards and Artificial Analysis and LMArena publications.",
      "Six tree rows added and one updated; the decision-model fine-tunes are reviewed and deferred pending the placement rules' 3-of-6 test rather than placed.",
      "Scorecard: Opus 5.5 takes the closed-frontier and reasoning leader slots from Fable 5.1 on the independent index; MiMo-V2.6-Pro takes the open leader slot from DeepSeek-V4.1-Flash on verified weights and an independent score; the Open frontier challenger moves from Atria Dawn Preview to GLM-5.3; the Coding leader and challenger move from Opus 5.5 and GPT-6 Astra to Sonnet 5.5 and GPT-6.1 Sol; the Edge / small leader moves from the Snapdragon 8 Elite Extreme on-device claim to the 4B llama.cpp decision heads and the challenger from Nex-N2.5 mini to Clef-flash; the Multimodal challenger moves from Gemini 3.8 Flash TTS to Gemini 4 Argon.",
      "Every benchmark figure is labelled vendor-run or independent at the point of use; vendor tables from different labs are not compared against each other except where flagged as such.",
      "Revision cycle 1 (editorial board): the Terminal-Bench 4.0 table now carries Artificial Analysis's independent runs for Sonnet 5.5 (64%), Opus 5.5 (60%), GPT-6 Astra (59%) and Argon (57%), which the first draft missed; Argon's hallucination figure is corrected from highest to lowest in the set, achieved by abstention; Argon's tier is corrected to frontier; the claim that OpenAI attributes Sol's cost advantage to cache reads is removed; the Opus 5.5 AutomationBench row is labelled derived; the bigRead is restructured around the three meters.",
      "Revision cycle 2 (editorial board): the Opus 5.5 AutomationBench row now carries Google's printed 42.5% instead of a derived figure and the Sol row is marked non-comparable; Argon's Omniscience comparison is restricted to the models in Artificial Analysis's September 30 article; the AA Terminal-Bench and Omniscience figures are labelled as outside the week's graded research set at the point of use (the fact audit verified them against the article; research.json still needs rows for them); the NOTICE-file provenance clause is removed as unverified; the Grok 4.7 context note no longer says 'undisclosed'; the Holo4 cost claim is narrowed to absolute per-task cost with no frontier comparison; the bigRead leads with the gating thesis, credits DataCamp, TokenCost and synthorai for the meter frame, and drops the 'house contribution' line; the mid-tier meter packet's fact home is the first architecture-watch entry and the AISI table's fact home is Agent Techniques; the Edge / small tier is defined and re-ranked; the Coding and Open frontier changes are recorded in the scorecard changelog line."
    ]
  }
}
