{
  "_meta": {
    "publication": "The Model Pulse",
    "schemaVersion": "2026.05.02",
    "generatedAt": "2026-09-01T23:37:09.863Z",
    "canonicalUrl": "https://brianletort.ai/industry/models/2026-W35",
    "markdownUrl": "https://brianletort.ai/industry/models/2026-W35/llm.md",
    "pdfUrl": "https://brianletort.ai/downloads/model-pulse-2026-W35.pdf",
    "treeUrl": "https://brianletort.ai/industry/tree",
    "sourceFile": "src/data/industry/models/2026-W35.ts"
  },
  "issue": {
    "slug": "2026-W35",
    "isoYear": 2026,
    "isoWeek": 35,
    "issueNumber": 19,
    "publishedAt": "2026-08-29",
    "cadence": "weekly",
    "periodLabel": "Week 35 of 2026",
    "bigRead": {
      "headline": "Kodam's practitioner harness result and seat metering, not Hot Chips silicon — the week's model procurement read",
      "body": "Vijay Kodam's Aug 24 practitioner bakeoff documents an 8× wall-clock swing on identical Qwen3.8-27B weights from the harness-plus-runtime pairing — Pi+Ollama 34 minutes versus Qwen Code+LM Studio 4h46m, with his own Pi+LM Studio control at 115.8 minutes, so neither the harness nor the inference engine alone explains the spread (grade-3 bakeoff in Agent Techniques proofOfValue[1]); GitHub Copilot promotional credits expire September 1 — see Application Layer pricingShifts for Business and Enterprise credit math. For near-term agent economics, test those practitioner-documented integration and metering effects before treating Hot Chips rack headlines as the binding lever — local wall-clock and datacenter tokens-per-megawatt are different axes, not a ranking. Managed-stack buyers should note OpenAI's GPT-5.6-in-Kiro reports ~82% cost reduction from tier routing alone (software-02).\n\nFive open-weight tree rows landed August 24–28 plus Ox Alpha stealth demand (community-reported, grade 3): GLM-5.3-Flash MIT weights, Qwen3.8-Flash-Next architecture preview, IBM Granite 4.2 at vendor-reported 57.00% SWE-Bench Verified (30B), Tencent Hy4 preview, and Thomson-1.0-Small as a Qwen3.6 derivative. OpenAI integrated GPT-5.6 Sol, Terra, and Luna into AWS Kiro with vendor-reported ~82% cost reduction per successful Terminal-Bench 2.1 task versus an unspecified baseline.\n\nQwen3.8-Flash-Next ships as an open checkpoint under qwen-community-1.0 while production API SKU Qwen3.8-Flash at $0.16/$0.47 per million tokens with 1M default context is a separate hosted product — procurement must not treat the 360GB artifact and the metered API as one SKU.\n\nIBM Granite 4.2 is the week's clearest enterprise on-prem license play: Apache 2.0, switchable chain-of-thought, native tool calling, and sandbox agentic RL yielding vendor-reported 57.00% SWE-Bench Verified at 30B. Read the terminal-agent row before generalizing that: IBM's own table puts the 30B at 29.24% on Terminal-Bench 2.1, roughly 55 points below GLM-5.3-Flash's vendor-reported 84.3% on the same benchmark, so Granite is a deployability and licensing choice rather than a capability one. Granite Speech 5.0 Turbo ASR on the same release day reports RTFx (real-time factor — throughput relative to audio duration) near 12,600 on a single H200 per IBM — making speech-to-agent loops viable without a hyperscaler API. Thomson Reuters spent ~$40M continual-learning open weights: the Thomson LLM that reaches customers is a closed deployment built from Alibaba's Qwen3.5-397B open weights, and Thomson-1.0-Small is an open-weight Qwen3.6-35B-A3B derivative — see Application Layer for vertical packaging on Claudeforce and Gemini Enterprise."
    },
    "treeDelta": {
      "summary": "5 model rows added and 1 updated. The canopy widened across open MoE (GLM-5.3-Flash, Hy4 preview), Qwen4 architecture preview (Qwen3.8-Flash-Next), enterprise Apache 2.0 agent stack (Granite 4.2), and legal vertical derivative (Thomson-1.0-Small). GLM-5.3 parent row updated to reflect Flash weight release.",
      "added": [
        "glm-5-3-flash",
        "qwen3-8-flash-next",
        "granite-4-2-30b",
        "thomson-1-0-small",
        "hy4-preview"
      ],
      "updated": [
        "glm-5-3"
      ],
      "note": "GLM-5.3-Flash confirms Ox Alpha, and Z.ai closed the weights promise days later: the full 753B GLM-5.3 checkpoint published to Hugging Face Aug 27–28 under a bespoke license (Z.AI security review required for Model-as-a-Service operators above $10B revenue), not MIT like Flash. Qwen3.8-Flash-Next is a preview checkpoint — production API is separate SKU Qwen3.8-Flash."
    },
    "frontierMovements": [
      {
        "modelId": "gpt-5-6",
        "name": "GPT-5.6 in AWS Kiro",
        "vendor": "OpenAI",
        "releaseDate": "2026-08-24",
        "headline": "Full GPT-5.6 family (Sol, Terra, Luna) inside AWS Kiro spec-driven coding agent — vendor-reported ~82% lower cost per successful Terminal-Bench 2.1 task versus unspecified baseline",
        "why": "Buyers routing agentic coding through AWS should evaluate Kiro as a bundled harness plus model tier choice, not as raw API pricing alone. OpenAI documents ~82% cost reduction per successful Terminal-Bench 2.1 task in Kiro on its blog; demand baseline definitions before moving production traffic.",
        "tier": "frontier",
        "architecture": "reasoning",
        "source": "OpenAI",
        "sourceUrl": "https://openai.com/index/gpt-5-6-in-kiro/"
      },
      {
        "modelId": "glm-5-3-flash",
        "name": "GLM-5.3-Flash (Ox Alpha reveal)",
        "vendor": "Z.AI (Zhipu)",
        "releaseDate": "2026-08-26",
        "headline": "320B/18B-active natively multimodal MoE MIT weights — six-day stealth run processed 20T+ OpenRouter tokens before identity reveal",
        "why": "Route high-volume coding and agent loops to GLM-5.3-Flash for pilot this quarter if cost dominates, but model post-promo pricing September 9 and independent AA runs after the free stealth period. Serving-hardware claims circulating this week are ungraded: cheap inference under a promotion is not evidence of export-control training independence.",
        "tier": "frontier",
        "architecture": "moe",
        "source": "Z.ai",
        "sourceUrl": "https://z.ai/blog/glm-5.3-flash"
      }
    ],
    "openWeights": [
      {
        "modelId": "glm-5-3-flash",
        "name": "GLM-5.3-Flash",
        "vendor": "Z.AI (Zhipu)",
        "releaseDate": "2026-08-26",
        "headline": "MIT-licensed 320B/18B-active multimodal MoE with vendor-reported 84.3% Terminal-Bench 2.1 and AA Intelligence Index 57 at $0.045/task",
        "why": "The artifact is real and downloadable; every headline score is vendor-run or measured during free anonymous access. Teams needing MIT-clean weights for agent pilots should download now; teams needing audited leaderboard position should wait for Artificial Analysis refresh post-promo.",
        "tier": "open_frontier",
        "architecture": "moe",
        "source": "Hugging Face",
        "sourceUrl": "https://huggingface.co/zai-org/GLM-5.3-Flash"
      },
      {
        "modelId": "qwen3-8-flash-next",
        "name": "Qwen3.8-Flash-Next",
        "vendor": "Alibaba",
        "releaseDate": "2026-08-26",
        "headline": "Qwen4 hybrid DeltaNet + Gated Attention preview — 125B/6B active, 512 experts, ~360GB FP checkpoint",
        "why": "Architecture preview for cost-efficient agent routing; self-host only if you can absorb 360GB and qwen-community-1.0 license terms. Most enterprises should start on hosted Qwen3.8-Flash API at $0.16/$0.47 with 1M context rather than this checkpoint.",
        "tier": "open_frontier",
        "architecture": "moe",
        "source": "Alibaba Cloud",
        "sourceUrl": "https://www.alibabacloud.com/blog/qwen3-8-flash-next-a-new-architecture-towards-ultimate-cost-efficiency_603501"
      },
      {
        "modelId": "granite-4-2-30b",
        "name": "IBM Granite 4.2 (30B)",
        "vendor": "IBM",
        "releaseDate": "2026-08-25",
        "headline": "Apache 2.0 agentic RL family — vendor-reported 57.00% SWE-Bench Verified and 29.24 Terminal-Bench 2.1 at 30B",
        "why": "Default open-weight pick for regulated on-prem agent estates this quarter. Pair with Granite Speech 5.0 Turbo for voice-to-agent loops without frontier API dependency. Require independent SWE-Bench reproduction before substituting for closed frontier on highest-risk code paths.",
        "tier": "open_frontier",
        "architecture": "reasoning",
        "source": "IBM Research",
        "sourceUrl": "https://research.ibm.com/blog/introducing-granite-4-2"
      },
      {
        "modelId": "hy4-preview",
        "name": "Tencent Hy4 Preview",
        "vendor": "Tencent",
        "releaseDate": "2026-08-28",
        "headline": "770B/49B-active MoE, 1M context, Apache 2.0 — internal blind eval 2.99/4.00 on 203 engineering tasks",
        "why": "Strongest Chinese open MoE drop of the week on license terms, but internal eval is vendor-run. Pilot on FP8 vLLM images if you already run Tencent stack; otherwise wait for independent coding leaderboard entries.",
        "tier": "open_frontier",
        "architecture": "moe",
        "source": "Tencent",
        "sourceUrl": "https://www.tencent.com/tencent-releases-and-open-sources-tencent-hy4-preview/"
      },
      {
        "modelId": "thomson-1-0-small",
        "name": "Thomson-1.0-Small",
        "vendor": "Thomson Reuters",
        "releaseDate": "2026-08-24",
        "headline": "Qwen3.6-35B-A3B derivative on Hugging Face under non-commercial academic license — not Apache open weights",
        "why": "Legal teams should treat this as a corpus-tuned derivative for research and non-commercial experimentation, not as a substitute for Thomson's closed-deployment Thomson LLM (itself a continual-learning derivative of Qwen3.5-397B open weights per the Thomson technical report) in production CoCounsel. Editorial forecast: lineage disclosure may become an RFP requirement — the falsifier (two top-tier vertical platforms train fully proprietary frontier models with no open-weight base disclosed, by Q2 2027) is stated once in the Weekly synthesis.",
        "tier": "specialist",
        "architecture": "moe",
        "source": "Thomson Reuters",
        "sourceUrl": "https://huggingface.co/thomsonreuters/Thomson-1.0-Small"
      }
    ],
    "architectureWatch": [
      {
        "pattern": "Open-weight releases bifurcate hosted API SKUs from downloadable checkpoints",
        "examples": [
          "Qwen3.8-Flash vs Flash-Next",
          "GLM-5.3-Flash stealth vs list API",
          "Thomson closed vs Small derivative"
        ],
        "body": "Three vendors shipped paired artifacts this week where the production metered endpoint differs from the Hugging Face row — different context defaults, tool interfaces, or licenses. Procurement contracts must name the SKU (hosted API id or checkpoint hash), not the marketing family name, because substitutability failed on every paired release.",
        "source": "Alibaba Cloud, Z.ai, Thomson Reuters",
        "sourceUrl": "https://www.alibabacloud.com/blog/qwen3-8-flash-next-a-new-architecture-towards-ultimate-cost-efficiency_603501"
      },
      {
        "pattern": "Agentic RL in open-weight post-training is becoming table stakes for coding MoE",
        "examples": [
          "IBM Granite 4.2 sandbox RL",
          "GLM-5.3-Flash terminal benchmarks",
          "Hy4 self-optimization"
        ],
        "body": "IBM's Granite 4.2 explicitly trains in sandbox software-engineering, terminal, and web-search environments; Z.ai and Tencent headline terminal and SWE scores on the same week's releases. The architecture shift is not bigger MoE alone — it is RL inside harness-shaped sandboxes. Buyers evaluating open weights for agents should require disclosure of sandbox design, not just parameter counts.",
        "source": "IBM Research, Z.ai, Tencent",
        "sourceUrl": "https://research.ibm.com/blog/introducing-granite-4-2"
      },
      {
        "pattern": "Vertical corpora plus orchestration atop semi-open foundations replaces from-scratch frontier training",
        "examples": [
          "Thomson on Qwen3.6",
          "CoCounsel Claude Agent SDK",
          "Gemini Enterprise verticals",
          "Fuse Terminal GA",
          "Ling 3.0 Flash Fin free window"
        ],
        "body": "Thomson Reuters spent ~$40M on continual learning over open weights: Thomson LLM is a closed deployment derived from Qwen3.5-397B and Thomson-1.0-Small an open Qwen3.6-35B-A3B derivative, while CoCounsel retains Claude Agent SDK for agentic workflows (Law.com, Aug 24). Fuse Terminal GA ($999/mo, never-gated API) and Ling 3.0 Flash Fin free through Sep 25 offer finance alternatives to packaged verticals. Regulated vertical moats consolidate on corpus plus platform-governed orchestration — procurement should ask foundation lineage questions in every vertical RFP.",
        "source": "Thomson Reuters, Law.com",
        "sourceUrl": "https://www.law.com/legaltechnews/2026/08/24/thomson-reuters-launches-proprietary-llm-thomson-updates-cocounsel/"
      }
    ],
    "benchmarkMoves": [
      {
        "benchmark": "Terminal-Bench 2.1 (vendor-reported coding agent rows)",
        "movement": "GLM-5.3-Flash publishes vendor-reported 84.3%; IBM Granite 4.2 30B reports vendor-reported 29.24% on the same resolve-rate scale as its 57.00% SWE-Bench Verified row in IBM's own results table — a roughly 55-point gap on the same benchmark, which is the honest read and the reason the open-frontier pick is a licensing decision, not a capability one; GPT-5.6 Terra vendor-reported 87.4% cited on Z.ai comparison table. Harness and tier choice still dominate cross-vendor comparability.",
        "rows": [
          {
            "model": "GPT-5.6 Terra (vendor table cite)",
            "score": "87.4%"
          },
          {
            "model": "GLM-5.3-Flash (vendor)",
            "score": "84.3%"
          },
          {
            "model": "IBM Granite 4.2 30B (vendor)",
            "score": "29.24%"
          }
        ],
        "source": "Z.ai, IBM Research",
        "sourceUrl": "https://z.ai/blog/glm-5.3-flash"
      },
      {
        "benchmark": "SWE-Bench Verified",
        "movement": "IBM Granite 4.2 30B reports 57.00% — strongest audited-style vendor claim from an Apache 2.0 open release this week; independent reproduction pending.",
        "rows": [
          {
            "model": "IBM Granite 4.2 30B (vendor)",
            "score": "57.00%"
          },
          {
            "model": "GLM-5.3-Flash (vendor DeepSWE)",
            "score": "63.4% DeepSWE v1.1"
          }
        ],
        "source": "IBM Research",
        "sourceUrl": "https://research.ibm.com/blog/introducing-granite-4-2"
      },
      {
        "benchmark": "Artificial Analysis Intelligence Index v4.1.1",
        "movement": "GLM-5.3-Flash scores 57 at vendor-reported $0.045/task during launch window; board unchanged for closed frontier leaders from August 6 refresh.",
        "rows": [
          {
            "model": "Claude Opus 5",
            "score": "63 (as of Aug 6)"
          },
          {
            "model": "GLM-5.3",
            "score": "60 (as of Aug 6)"
          },
          {
            "model": "GLM-5.3-Flash",
            "score": "57 at $0.045/task (vendor)"
          }
        ],
        "source": "Artificial Analysis, Z.ai",
        "sourceUrl": "https://artificialanalysis.ai/models/glm-5-3-flash/providers"
      },
      {
        "benchmark": "Open ASR Leaderboard throughput (vendor-reported)",
        "movement": "IBM Granite Speech 5.0 Turbo CTC reports vendor-reported RTFx (real-time factor — audio processed faster than playback) ~12,600 on single H200 — ~2× leaderboard leaders per IBM.",
        "rows": [
          {
            "model": "Granite Speech 5.0 Turbo (vendor)",
            "score": "RTFx ~12,600"
          },
          {
            "model": "Leaderboard leaders (IBM cite)",
            "score": "RTFx ~6,000"
          }
        ],
        "source": "IBM Research",
        "sourceUrl": "https://research.ibm.com/blog/introducing-granite-4-2"
      }
    ],
    "scorecard": {
      "asOf": "2026-08-29",
      "rows": [
        {
          "tier": "Closed frontier",
          "leader": "Claude Opus 5",
          "challenger": "GPT-5.6 Sol Max",
          "note": "Capability board unchanged from August 6 AA refresh. NVIDIA and OpenAI reframed economics toward agentic tokens-per-megawatt and Kiro tier routing — decide on cost per completed agent task, not list token price alone."
        },
        {
          "tier": "Open frontier",
          "leader": "GLM-5.3-Flash",
          "challenger": "IBM Granite 4.2 30B",
          "note": "GLM-5.3-Flash wins on vendor-reported terminal scores and MIT license; Granite wins on Apache 2.0 enterprise deployability and SWE-Bench Verified headline. Neither has independent tracker confirmation at publish — pick license and hosting boundary first, capability second."
        },
        {
          "tier": "Reasoning",
          "leader": "Claude Opus 5",
          "challenger": "GPT-5.6 Terra",
          "note": "Terra enters as the efficiency tier inside Kiro with vendor-reported 87.4% Terminal-Bench 2.1 on Z.ai comparison table — route spec-driven coding to Terra when vendor controls harness; Opus remains default for highest-stakes reasoning until independent agentic cost curves land."
        },
        {
          "tier": "Coding",
          "leader": "Claude Fable 5 Max",
          "challenger": "GLM-5.3-Flash",
          "note": "GLM-5.3-Flash vendor-reported 84.3% Terminal-Bench 2.1 is not harness-comparable to Fable rows without independent reproduction. Local open-weight coding default shifts to Granite 4.2 30B for Apache estates and GLM-5.3-Flash for MIT high-volume pilots."
        },
        {
          "tier": "Multimodal",
          "leader": "Gemini 3.7 Flash",
          "challenger": "GLM-5.3-Flash",
          "note": "GLM-5.3-Flash adds natively multimodal MIT MoE at datacenter scale; Gemini retains hosted enterprise vertical packaging. Qwen3.8-Flash-Next preview adds architecture option for self-host multimodal MoE at extreme footprint."
        },
        {
          "tier": "Edge / small",
          "leader": "Granite Speech 5.0 Turbo",
          "challenger": "IBM Granite 4.2 8B",
          "note": "Granite Speech 5.0 Turbo vendor-reported RTFx ~12,600 on H200 is the week's underplayed enterprise angle for voice-to-agent loops. Granite 4.2 3B/8B dense with tool calling challenges edge leader on enterprise agent features."
        }
      ]
    },
    "vendorSignals": [
      {
        "vendor": "Z.ai",
        "date": "2026-08-26",
        "signal": "Retired Ox Alpha alias and published GLM-5.3-Flash MIT weights after 20T+ token stealth run; list API $0.15/$0.50 with 50% promo through Sep 9",
        "meaning": "Stealth free access was a demand test, not steady-state pricing. Renegotiate agent contracts before September 9 promo expiry and require post-promo rate cards in writing.",
        "source": "Z.ai",
        "sourceUrl": "https://z.ai/blog/glm-5.3-flash"
      },
      {
        "vendor": "Alibaba",
        "date": "2026-08-26",
        "signal": "Qwen3.8-Flash production API at $0.16/$0.47 per million tokens with 1M default context; Flash-Next open checkpoint separate SKU",
        "meaning": "Hosted Qwen agent routing at sub-frontier list price with million-token default — pilot BYOK agent loops on Qwen3.8-Flash API before reserving Rubin capacity if harness overhead dominates wall-clock.",
        "source": "Alibaba Cloud",
        "sourceUrl": "https://www.alibabacloud.com/blog/qwen3-8-flash-next-a-new-architecture-towards-ultimate-cost-efficiency_603501"
      },
      {
        "vendor": "OpenAI",
        "date": "2026-08-24",
        "signal": "GPT-5.6 family integrated into AWS Kiro with Terra positioned as Terminal-Bench efficiency tier",
        "meaning": "AWS buyers get vendor-controlled harness plus model tier in one procurement line — shifts evaluation from raw API to integrated agent workflow cost.",
        "source": "OpenAI",
        "sourceUrl": "https://openai.com/index/gpt-5-6-in-kiro/"
      },
      {
        "vendor": "IBM",
        "date": "2026-08-25",
        "signal": "Granite 4.2 Apache 2.0 with agentic RL; Granite Speech 5.0 Turbo ASR at vendor-reported ~12,600 RTFx",
        "meaning": "IBM re-enters open-weight agent stack for regulated on-prem — competitive against hosted copilots for enterprises blocked from hyperscaler APIs.",
        "source": "IBM Research",
        "sourceUrl": "https://research.ibm.com/blog/introducing-granite-4-2"
      },
      {
        "vendor": "GitHub",
        "date": "2026-08-28",
        "signal": "Promotional Copilot credit pools expire Sep 1 — see Application Layer pricingShifts for Business 3,000→1,900 and Enterprise 7,000→3,900 credit math",
        "meaning": "Metering contraction on the largest developer surface — teams built on promo allowances face overage unless admins cap spend or route to BYOK open-weight stacks validated this week.",
        "source": "GitHub Changelog",
        "sourceUrl": "https://github.blog/changelog/2026-08-28-upcoming-changes-to-github-copilot-policies-and-billing/"
      }
    ],
    "watchlist": [
      {
        "window": "Sep 3",
        "title": "FERC (Federal Energy Regulatory Commission) comment deadline on PJM Interim Resource Adequacy Service (ER26-3515)",
        "why": "≥50 MW loads without bring-your-own-capacity may face curtailment-before-residential rules — affects Virginia and PJM hyperscale energization math."
      },
      {
        "window": "Sep 9",
        "title": "GLM-5.3-Flash launch promo expires",
        "why": "Steady-state API pricing resets — rerun agent business cases on list $0.15/$0.50 and watch Artificial Analysis provider variance."
      },
      {
        "window": "Sep 15",
        "title": "GLM-5.3 license review gate (prediction p90 resolved hit — 753B weights on Hugging Face Aug 27–28)",
        "why": "Full weights shipped under a bespoke non-MIT license — Model-as-a-Service operators above $10B revenue need Z.AI security review before commercial use; watch whether that gate changes provider availability on OpenRouter and Artificial Analysis."
      },
      {
        "window": "Sep 25",
        "title": "Ling 3.0 Flash Fin free API window ends",
        "why": "Finance-tuned MoE free hosted access through OpenRouter/Vercel Gateway — last day to benchmark spreadsheet and disclosure-retrieval agents at zero list cost."
      },
      {
        "window": "Oct 2026",
        "title": "MetaRoCE OCP specification at Global Summit",
        "why": "Loss-tolerant RDMA transport spec drop — sets whether million-GPU Ethernet stays merchant-open or NVIDIA-codesigned Multiplane wins by default."
      },
      {
        "window": "Q3 FY2027 print",
        "title": "NVIDIA Vera Rubin mix disclosure in earnings",
        "why": "CFO guided ~20% datacenter revenue from Rubin in Q3 — separates sell-in allocation from fleet utilization on agentic workloads."
      }
    ],
    "changelog": [
      "Five tree rows added (GLM-5.3-Flash, Qwen3.8-Flash-Next, Granite 4.2, Thomson-1.0-Small, Hy4 preview); GLM-5.3 updated for Flash weight release.",
      "Big read leads with practitioner harness variance (Kodam Aug 24) and Sep 1 Copilot metering over Hot Chips silicon — SKU-split procurement read is the differentiated contribution.",
      "Scorecard open frontier row shifts to GLM-5.3-Flash vs Granite 4.2 on license and deployability, not leaderboard position alone.",
      "Revision 3: Granite Terminal-Bench 2.1 restated as 29.24% (same resolve-rate scale as IBM's SWE-Bench row, not an 'IBM scale'), recommendation reframed as a licensing pick with the ~55-point gap to GLM-5.3-Flash stated; Kodam bakeoff reframed around the harness-runtime pairing (Pi+LM Studio control cited) and the 8×-versus-rack-efficiency ranking removed; Qwen3.8-Flash input price corrected to $0.16; Thomson lineage corrected (Thomson LLM is a Qwen3.5-397B derivative, Small a Qwen3.6 derivative); GLM-5.3 note and p90 watchlist updated for the Aug 27–28 full-weight release; unsupported domestic-accelerator claim removed; Thomson falsifier aligned with the weekly polarity."
    ]
  }
}
