brianletort.ai
All issues

The Model Pulse

Issue 16 · Week 32 of 2026.

/Weekly read/~6 min read/Public sources onlyDownload brief

The Big Read

Two new frontier models shipped and the more important releases were the measuring instruments — the same weights now score differently depending on who serves them

The thesis this issue defends

Artificial Analysis launched the Endpoint Accuracy Index on Aug 4 and it reframes every open-weights procurement decision this publication has made. The method is explicit: benchmark each serverless endpoint against a self-hosted reference deployment of the official weights at the lab's recommended precision, across tool calling (BFCL-500), scientific reasoning (HLE-250) and long-context recall (AA-LCR-25). The findings are not marginal. Some gpt-oss-120b endpoints score 22% on BFCL-500 against 37% for the reference, driven by tool-call parsing and formatting differences. The most restrictive GLM-5.2 endpoints score half the reference or less on HLE-250 because output-token limits truncate reasoning before it completes. DeepSeek V4 Pro is the good case — most endpoints at parity, with DeepSeek's own first-party endpoint slightly above reference. Endpoints scoring below reference generally produce fewer output tokens per task, and the worst produce about half. A model ID is no longer sufficient to specify what you bought.

Two days later the same firm published Intelligence Index v4.1.1, a methodology patch that moved HLE, AA-LCR and AA-Omniscience grading to GPT-5.6 Luna, unifying three evaluations under one modern grader. Most models moved under a point. The largest mover was Muse Spark 1.2 at +2.7 — almost as much as the model gained from its own release (51 to 54 from Muse Spark 1.1). Pin the Index version in any comparison you carry forward, because a grader change is now capable of moving a model nearly as much as a generation.

The actual releases were Meta's Muse Spark 1.2 (Aug 5) and Alibaba's Qwen3.8-Max API GA (Aug 3), and both tell a cost story rather than a capability story. Muse Spark 1.2 rose to Index 54 pre-patch and posted a GDPval-AA v2 Elo of 1631, up 260 points from 1.1, at unchanged list pricing of $1.25/$4.25 — yet Artificial Analysis measured cost per Index task rising from $0.29 to $0.40, roughly 38%, because input tokens rose ~53% and output tokens ~36% per task. Qwen3.8-Max landed at Index 58 and consumed 150M output tokens to run the Index against a class median of 66M, which AA labels 'very verbose'. Rate cards did not move; bills did.

Two honesty notes. Muse Spark 1.2's AA-Omniscience gain (18 to 22) comes with hallucination falling from 38% to 28% while the attempt rate fell from 82% to 67% and accuracy fell from 41% to 38% — the model is refusing more, not knowing more, and the index does not penalize refusal. And Liquid's LFM2.5-2.6B shipped on Aug 4 with a release page describing deployment 'without restrictions' while the LICENSE in the same repository withholds commercial use from any entity above $10M in revenue. Net/net: pin the endpoint, pin the Index version, read the license, and measure cost per completed task rather than per token.

Tree delta

What changed in the tree.

3 models added, 1 updated.

Muse Spark 1.2 and Qwen3.8-Max join as closed frontier rows, and LFM2.5-2.6B joins the edge branch as the first entry recorded with a license-contradiction annotation rather than a clean open-weight classification. Kimi K3 updates to carry Artificial Analysis's new 'Open Weights (Commercial Use Restricted)' category label, which the firm introduced this week because revenue-gated open weights have become common enough to institutionalize.

Added (3)

  • muse-spark-1-2
  • qwen3-8-max
  • lfm2-5-2-6b

Updated (1)

  • kimi-k3

Qwen3.8-Max enters with a disclosed conflict rather than a settled spec: Alibaba's launch blog describes 2.4T MoE with 95B active parameters while Artificial Analysis's model page states Alibaba has not disclosed model size or parameter count. The row records the vendor figures as vendor-stated and uncorroborated. Open weights for the Qwen-Max class and a companion Qwen3.8-27B were announced as coming 'next week' and had not shipped as of Aug 8, so neither is added — the W30 Kimi K3 lesson applies. OpenAI's Astra is deliberately excluded: it has not shipped, and its only public capability signal is a safety flag rather than a benchmark.

Explore the LLM Evolutionary Tree

Frontier movements

Flagship-class releases.

4 releases this period.

Vendor-stated frontier capability. The releases that reset the closed-source ceiling.

  • /Meta/Frontier/Reasoning

    Muse Spark 1.2

    Index 54 pre-patch (+3 over 1.1) and GDPval-AA v2 Elo 1631, up 260 points — at a 38% higher cost per task on unchanged list pricing

    Re-baseline on cost per completed task, not the rate card: pricing held at $1.25/$4.25 per 1M with $0.15 cached input, but Artificial Analysis measured cost per Index task rising from $0.29 to $0.40 because input tokens rose ~53% and output ~36% per task, concentrated in GDPval. Terminal-Bench v2.1 moved 78% to 80% and τ³-Banking 25% to 27%, while SciCode fell 58% to 56% and HLE 45% to 44%. Meta gave AA pre-release access, and Meta's own coding comparison ran 1.1 in mini-swe-agent against 1.2 in Muse Code, so part of the vendor-reported coding delta is harness rather than model.

    Meta AI Research; Artificial Analysis (independent); Meta developer pricing docs

  • /Alibaba / Qwen/Frontier/MoE

    Qwen3.8-Max

    API GA at Index 58 with 1M context and $2.00/$6.00 pricing — and 150M output tokens to run the Index against a 66M class median

    Treat verbosity as the procurement risk here: AA labels it 'very verbose' and spent $1,741.41 evaluating it on the Index, so blended token assumptions built from other models will understate the bill. Cache hits at $0.25 cut input cost 88%, which makes prompt-caching discipline unusually valuable on this model. Reasoning model with text, image and video input; blended 7:2:1 rate is $1.18 per 1M. Vendor claims 2.4T MoE with 95B active parameters, but AA states Alibaba has not disclosed size — treat parameters as uncorroborated.

    Artificial Analysis model page (independent); Alibaba Cloud launch blog (vendor)

  • /OpenAI/Reasoning/Agentic

    Astra (unreleased)

    Launch delayed after internal evaluations could not rule out a Critical cybersecurity capability level — the first such flag OpenAI has made

    Do not put Astra in a roadmap slot yet. OpenAI reported 'significant advancements in agentic coding and cybersecurity' strong enough that it cannot rule out Critical under its Preparedness Framework, paused internal activities not meeting stricter controls, and Altman confirmed the delay. No benchmark numbers, no ship date, and no final rating exist. The Framework's stated policy at Critical is to halt development, and OpenAI paused selectively rather than halting — which implies its own assessment is that the threshold is not yet met.

    OpenAI; The Decoder

  • /OpenAI/Frontier/Reasoning

    GPT-5.6 Sol / Luna ChatGPT surface update

    Product-surface change only: ChatGPT chat moves to August Sol/Luna variants while Codex and ChatGPT Work stay on the July checkpoints

    This is a naming hazard, not a capability event. The same model name now refers to different checkpoints depending on which surface serves it, which breaks the assumption that an internal evaluation of 'GPT-5.6 Sol' transfers across products. No API checkpoint changed and no API price changed. Safety designation is unchanged at High for Cyber and Bio/Chem. Vendor-internal factual-error reduction claims circulated via secondary press; treat as vendor-stated.

    OpenAI Deployment Safety Hub; secondary press for the internal statistics

Open weights

Open-frontier and open-source drops.

2 releases this period.

Open-weights releases that change procurement options. Pull these into pilot when score parity meets license parity.

  • /Liquid AI/Edge / small/Dense

    LFM2.5-2.6B

    2.6B on ~34T tokens with 128K context and strong on-device throughput — under a license withholding commercial use above $10M revenue

    Send this to counsel before any pilot. The release page says 'Open-weight — Download, fine-tune, and deploy without restrictions'; Section 5(b) of the LFM Open License v1.0 in the same repository says commercial use by an entity exceeding the $10M annual revenue Threshold is 'not licensed under this Agreement', Commercial Use reaches internal enterprise use, Legal Entity includes the controlled group, and Section 11 terminates automatically on non-compliance. This is materially stricter than Kimi K3's W31 terms, which gated hosting above $20M and pointed to a separate agreement — the LFM text names no such path. On capability: base and post-trained checkpoints, four-stage post-training ending in agentic RL inside real harnesses, 220 tok/s decode on an M5 Max and roughly 30 tok/s on a phone under 2.5 GB.

    Liquid AI blog; Hugging Face model card; LFM Open License v1.0 text

  • /Alibaba / Qwen/Open frontier/MoE

    Qwen-Max class open weights and Qwen3.8-27B (announced, not shipped)

    Announced as coming 'next week' alongside the Qwen3.8-Max API launch — no artifacts and no license text as of Aug 8

    Do not build a self-host plan on this. Nothing shipped in-window, the license is undisclosed, and Artificial Analysis's FAQ states plainly that Qwen3.8 Max is not open source. Track it as a promise with a deadline rather than an available artifact, and note that the license terms matter more than the weights given how the last three revenue-gated releases have read.

    Alibaba Cloud launch blog; Artificial Analysis model page FAQ

Architecture watch

Patterns to track.

4 patterns reshaping the canopy.

Architectural patterns that crossed multiple vendors this period. Each pattern lists exemplar releases and what it changes for deployment, cost, or capability.

  • The serving endpoint is now an architecture variable, not a delivery detail

    gpt-oss-120b: some endpoints at 22% on BFCL-500 against a 37% self-hosted referenceGLM-5.2: most restrictive endpoints at half the reference or less on HLE-250 from output-token capsDeepSeek V4 Pro: majority of endpoints at parity, first-party endpoint slightly above reference

    Artificial Analysis identified two concrete mechanisms rather than asserting variance. Tool-call parsing and formatting differ by provider, which is what moves BFCL-500. Output-token limits truncate reasoning before it completes, which is what moves HLE-250 — and the firm notes that endpoints scoring below reference generally produce fewer output tokens per task, with the worst producing roughly half. For architects the consequence is that an inference provider is now part of the model's effective configuration, and any evaluation run against a different endpoint than production is measuring something else. Add output-token limits, precision, and tool-call format to the acceptance checklist alongside the model ID.

    Artificial Analysis

  • Capability is rising while cost per task rises with it, at unchanged rate cards

    Muse Spark 1.2: $0.29 to $0.40 per Index task, input tokens +53% and output +36%Qwen3.8-Max: 150M output tokens to run the Index against a 66M class medianComparison set: Grok 4.5 (high) $0.37, GPT-5.6 Sol (medium) $0.39, Kimi K3 (max) $0.86

    Two independent releases this week got better and more expensive per unit of work without changing a published price. The mechanism is token consumption: reasoning models that think longer produce more billable output, and the effect concentrates in the long-horizon agentic evaluations where the capability gains also show up. This inverts the routing heuristic that has held for most of the year, where a cheaper rate card reliably meant a cheaper workload. Teams should re-derive per-task cost on their own traffic after every model upgrade, because the rate card no longer carries that information.

    Artificial Analysis

  • Abstention is being rewarded by hallucination metrics that do not price refusal

    Muse Spark 1.2 AA-Omniscience Index 18 to 22Hallucination rate 38% to 28%, but attempt rate 82% to 67% and accuracy 41% to 38%

    The headline reading of Muse Spark 1.2 is a 10-point hallucination reduction. The mechanism is that the model attempts 15 points fewer questions and its accuracy on what it does attempt fell three points. AA-Omniscience does not penalize declining to answer, so a more cautious model scores better without knowing more. This is not a criticism of the benchmark, which is measuring what it says it measures, but it is a warning about how the number travels. Any enterprise using hallucination-rate improvements to justify deployment in a high-stakes workflow should ask for the attempt rate alongside it, because a model that refuses more is a different operational product than a model that is more accurate.

    Artificial Analysis

  • Revenue-gated open weights have institutionalized into their own category

    Artificial Analysis added an 'Open Weights (Commercial Use Restricted)' chart categoryLFM Open License v1.0: no commercial use above $10M revenue, automatic termination on breachKimi K3 (W31): commercial hosting above $20M trailing revenue requires a separate agreement

    The most telling artifact this week is not a license but a chart label. Artificial Analysis introduced a distinct category defined as weights available with commercial use limited, typically requiring a paid license — which means the pattern is now frequent enough that an independent tracker needed a bucket for it. The practical consequence for architecture is that 'open weights' has stopped being a single procurement path and has become at least three: unrestricted permissive licenses, revenue-gated licenses with a negotiation path, and revenue-gated licenses without one. Liquid's is the third kind, which is the most restrictive form yet seen at this scale.

    Artificial Analysis; LFM Open License v1.0

Benchmark moves

Where the leaderboard moved.

4 benchmarks shifted.

Benchmark deltas that change a procurement read. Scores reflect public leaderboards or vendor model cards as of publication.

  • Endpoint Accuracy Index (new)

    First independent measurement of how much reference accuracy each serverless endpoint preserves for identical open weights

    • gpt-oss-120b, worst endpoint vs reference (BFCL-500)22% vs 37%
    • GLM-5.2, most restrictive endpoints (HLE-250)≤50% of reference
    • DeepSeek V4 Pro, majority of endpointsAt reference parity
    • DeepSeek first-party endpointSlightly above reference

    Artificial Analysis

  • Artificial Analysis Intelligence Index v4.1.1

    Grader models unified under GPT-5.6 Luna for HLE, AA-LCR and AA-Omniscience; most models moved under a point

    • Claude Opus 5 (max)63, unchanged at #1
    • Muse Spark 1.2 (xhigh)+2.7, largest move of any model
    • Typical model movementUnder 1 point

    Artificial Analysis

  • Cost per Intelligence Index task

    Muse Spark 1.2 rose ~38% over 1.1 at unchanged list pricing, placing it mid-pack rather than cheap

    • Muse Spark 1.1$0.29
    • Grok 4.5 (high)$0.37
    • GPT-5.6 Sol (medium)$0.39
    • Muse Spark 1.2$0.40
    • Kimi K3 (max)$0.86

    Artificial Analysis

  • AA-Omniscience (Muse Spark 1.1 to 1.2)

    Index improved four points, but the gain is driven by the model attempting fewer questions rather than answering better

    • AA-Omniscience Index18 to 22
    • Hallucination rate38% to 28%
    • Attempt rate82% to 67%
    • Accuracy on attempted41% to 38%

    Artificial Analysis

Tier scorecard

Who leads, who pushes.

6 tiers · leaders as of Aug 8, 2026.

A snapshot of leader-vs-challenger by tier. Useful for procurement shortlists when matching workload to model class. Pair with the benchmark moves above for the underlying scores.

  • Closed frontier

    Leader: Claude Opus 5 — Index 63

    Challenger: GPT-5.6 Sol, Claude Fable 5, Kimi K3

    Unchanged at the top through the v4.1.1 patch; no closed frontier flagship shipped in-window and OpenAI delayed Astra.

  • Open frontier

    Leader: Kimi K3 — Index 57, revenue-tiered license

    Challenger: DeepSeek V4 Pro, GLM-5.2

    No new open frontier weights shipped. The Endpoint Accuracy Index now makes delivered capability depend on the serving provider, so a single leader score understates the spread.

  • Reasoning

    Leader: Claude Opus 5

    Challenger: Qwen3.8-Max — Index 58, 1M context

    Qwen3.8-Max is the week's strongest new reasoning entrant on score, and the most verbose on tokens consumed per task.

  • Coding

    Leader: Claude Fable 5

    Challenger: GPT-5.6 Sol, Muse Spark 1.2

    Epoch's MirrorCode leaderboard puts Fable 5 at 64% against GPT-5.6 Sol at 20% on long-horizon tasks; treat as a leaderboard snapshot rather than a published study.

  • Multimodal

    Leader: Qwen3.8-Max — text, image and video input

    Challenger: Muse Spark 1.2

    No dedicated multimodal flagship shipped in-window; the movement is frontier reasoning models absorbing video input rather than specialist models advancing.

  • Edge / small

    Leader: LFM2.5-2.6B — under 2.5 GB, ~30 tok/s on a phone

    Challenger: Qwen3.5-4B, gemma-4-E2B-it

    Best on-device throughput result of the week, and the only leader on this board whose license withholds commercial use from most enterprises.

Vendor signals

Pricing, gating, deprecation.

5 non-release signals worth tracking.

The non-release moves that shift vendor risk — pricing, deprecations, gating decisions, license changes — with a one-line procurement read.

  • /Meta

    Contributor tier prices Muse Spark 1.2 at $0.10/$0.20 against the standard $1.25/$4.25, in exchange for permission to train on your prompts and completions

    This is the clearest price yet put on enterprise data in a frontier API: roughly 92% off input and 95% off output for training rights, with rate limits applied per team rather than per key. Any team evaluating the cheap tier needs a data-governance decision before a cost decision, and the two tiers should never be mixed inside one gateway route without an explicit policy control.

    Meta developer pricing and rate limits documentation

  • /OpenAI

    Astra development partially paused and launch delayed after evaluations could not rule out a Critical cybersecurity capability level

    A published capability framework has now imposed a shipping delay at the leading lab, which converts safety frameworks from disclosure documents into schedule risk that customers must model. Any roadmap contingent on a named unreleased frontier model now carries a failure mode that is neither technical nor commercial.

    OpenAI; The Decoder

  • /OpenAI

    The same model name now maps to different checkpoints by surface — ChatGPT chat on August Sol/Luna variants, Codex and ChatGPT Work on the July ones

    Internal evaluations tied to a model name no longer transfer across products from the same vendor. Governance teams that approved 'GPT-5.6 Sol' for a use case need to record which surface and which checkpoint was tested, or the approval record is ambiguous.

    OpenAI Deployment Safety Hub

  • /Liquid AI

    Release page markets LFM2.5-2.6B as deployable 'without restrictions' while the LICENSE in the same repository withholds commercial use above $10M revenue

    Procurement cannot rely on vendor release pages for license classification, because this contradiction shipped inside a single artifact on a single day. Add a step that diffs the marketing claim against the LICENSE file before any open-weight model enters evaluation.

    Liquid AI blog; LFM Open License v1.0

  • /Artificial Analysis

    Open-weights charts now carry a distinct 'Open Weights (Commercial Use Restricted)' category for weights whose commercial use requires a paid license

    An independent tracker creating a permanent category is the strongest available evidence that revenue-gated open weights are a durable market structure rather than a run of individual vendor choices. Model shortlists should carry license class as a first-class field alongside score and price.

    Artificial Analysis

Watchlist

On the radar next.

5 catalysts to watch, starting By Aug 31.

Specific model-side catalysts in the next 7–30 days that would change the read materially. Watching these tells us whether the canopy is widening or thinning.

  • By Aug 31

    Astra's final Preparedness Framework rating and revised launch date

    OpenAI reported that it cannot rule out Critical, which is not a rating. The final assessment decides whether this is the first genuine top-level capability flag or a delay narrated in safety language.

  • By Aug 15

    Qwen-Max class open weights and the Qwen3.8-27B companion

    Announced as coming 'next week' on Aug 3 with no license disclosed. Given three consecutive revenue-gated open releases, the license text will matter more than the weights.

  • By Sept 30

    Endpoint Accuracy Index coverage expanding beyond the launch set

    Artificial Analysis named Kimi K3 as next. Broader coverage is what turns this from an interesting finding into a usable procurement instrument for open-weight shortlists.

  • By Oct 31

    The first Gemini release under Kavukcuoglu

    The flagship model is unreleased against a planned June launch and the team was reorganized on Aug 5. The next release is the first evidence of whether separating research from product execution changed shipping cadence.

  • Ongoing

    Whether cost-per-task figures get republished under Index v4.1.1

    The cost-per-task comparison set was published under the prior methodology one day before the patch. If those figures are not restated, the most useful economic comparison in the industry is versioned against a superseded index.

Edits this issue

  • Muse Spark 1.2, Qwen3.8-Max and LFM2.5-2.6B added to the LLM Evolutionary Tree; Kimi K3 updated to carry the new 'Open Weights (Commercial Use Restricted)' category label.
  • Qwen3.8-Max is recorded with a disclosed parameter-count conflict: Alibaba states 2.4T MoE with 95B active, while Artificial Analysis states no size has been disclosed. Vendor figures are marked uncorroborated rather than adopted.
  • LFM2.5-2.6B is the first tree entry annotated for a contradiction between its release page and its license rather than classified simply as open weights.
  • Astra is deliberately excluded from the tree: no weights, no benchmarks, no ship date, and a capability signal that is a safety flag rather than a measurement.
  • The scorecard's open frontier row now carries a caveat that a single leader score understates delivered capability, following the Endpoint Accuracy Index finding that serving provider materially changes measured accuracy.

About The Model Pulse

A weekly read on the software side of the AI stack. Anchored to the LLM Evolutionary Tree, which the brief annotates each week. The cross-stack flywheel (capital, hardware, networking) is covered in The AI Stack Weekly.

Authorship and sources

Compiled from public model cards, vendor blogs, leaderboards, and official lab announcements. Written by Brian Letort. Independent analysis. Not investment guidance.

Operate. Publish. Teach.