Artificial Analysis launched the Endpoint Accuracy Index on Aug 4 and it reframes every open-weights procurement decision this publication has made. The method is explicit: benchmark each serverless endpoint against a self-hosted reference deployment of the official weights at the lab's recommended precision, across tool calling (BFCL-500), scientific reasoning (HLE-250) and long-context recall (AA-LCR-25). The findings are not marginal. Some gpt-oss-120b endpoints score 22% on BFCL-500 against 37% for the reference, driven by tool-call parsing and formatting differences. The most restrictive GLM-5.2 endpoints score half the reference or less on HLE-250 because output-token limits truncate reasoning before it completes. DeepSeek V4 Pro is the good case — most endpoints at parity, with DeepSeek's own first-party endpoint slightly above reference. Endpoints scoring below reference generally produce fewer output tokens per task, and the worst produce about half. A model ID is no longer sufficient to specify what you bought.
Two days later the same firm published Intelligence Index v4.1.1, a methodology patch that moved HLE, AA-LCR and AA-Omniscience grading to GPT-5.6 Luna, unifying three evaluations under one modern grader. Most models moved under a point. The largest mover was Muse Spark 1.2 at +2.7 — almost as much as the model gained from its own release (51 to 54 from Muse Spark 1.1). Pin the Index version in any comparison you carry forward, because a grader change is now capable of moving a model nearly as much as a generation.
The actual releases were Meta's Muse Spark 1.2 (Aug 5) and Alibaba's Qwen3.8-Max API GA (Aug 3), and both tell a cost story rather than a capability story. Muse Spark 1.2 rose to Index 54 pre-patch and posted a GDPval-AA v2 Elo of 1631, up 260 points from 1.1, at unchanged list pricing of $1.25/$4.25 — yet Artificial Analysis measured cost per Index task rising from $0.29 to $0.40, roughly 38%, because input tokens rose ~53% and output tokens ~36% per task. Qwen3.8-Max landed at Index 58 and consumed 150M output tokens to run the Index against a class median of 66M, which AA labels 'very verbose'. Rate cards did not move; bills did.
Two honesty notes. Muse Spark 1.2's AA-Omniscience gain (18 to 22) comes with hallucination falling from 38% to 28% while the attempt rate fell from 82% to 67% and accuracy fell from 41% to 38% — the model is refusing more, not knowing more, and the index does not penalize refusal. And Liquid's LFM2.5-2.6B shipped on Aug 4 with a release page describing deployment 'without restrictions' while the LICENSE in the same repository withholds commercial use from any entity above $10M in revenue. Net/net: pin the endpoint, pin the Index version, read the license, and measure cost per completed task rather than per token.