For architects tracking model capability shifts.
April rewrote the floor: open weights crossed the closed frontier, and the safety-gated frontier became a procurement criterion.
April 2026 Recap · April 25, 2026
Big read
April 2026 was the most productive month in frontier model history. Eleven model rows landed in the LLM Evolutionary Tree across eight vendors, including the first open-weights MoE to score frontier-class on SWE-Bench Pro coding (Kimi K2.6 at 58.6 and GLM-5.1 at 58.4, both above GPT-5.4 at 57.7 and Claude Opus 4.6 at 57.3). DeepSeek V4 shipped in two configurations under MIT license with 1M-token context and 1.6T total parameters, moving on-prem coding from 'can we?' to 'which workload first?' Anthropic withheld its Mythos flagship on cyber-capability grounds after UK AISI confirmed autonomous offensive capability — an inflection that turns capability gating from a research-org concern into a procurement diligence requirement. The closed frontier responded: GPT-5.5 (Apr 23) and Claude Opus 4.7 (Apr 16) reset the closed-source ceiling, with adaptive-thinking and ultra-long-context as the new defaults. Read together, April was the month the canopy widened on three axes at once: open-vs-closed parity on coding, reasoning-as-default at every tier, and capability gating as a market signal.
Tree delta
11 model rows added to the tree in April, spanning frontier closed releases, open-frontier MoEs, gated safety-class previews, and reasoning-tier successors. The decoder-only zone widened most; the multimodal and reasoning branches absorbed every major release.
Registry movement
The Anthropic line gained two flagship releases plus one gated preview in a single month. DeepSeek shipped V4 in two scales (Pro and Flash) plus a same-month R2 reasoning successor — a cadence no other vendor matched in April.
- Added
- claude-opus-4-6, qwen-3-6-plus, claude-mythos, glm-5-1, claude-opus-4-7, qwen-3-6-35b-a3b, kimi-k2-6, gpt-5-5, deepseek-v4-pro, deepseek-v4-flash, deepseek-r2
- Updated
- claude-opus-4-5, kimi-k2-5, gpt-5-4, qwen-3-5
Frontier movements
Anthropic · 2026-04-16 · Frontier · Agentic
Claude Opus 4.7
Reset the closed-frontier reasoning ceiling. The agentic surface area (computer-use, MCP, tool-use) makes Opus 4.7 the default reference for closed-frontier procurement when adaptive thinking and long-horizon agent runs are required. Pair with Sonnet 4.6 for the cost-tiered deployment.
- Model registry ID
- claude-opus-4-7
Sources anthropic.com release notes, model card
OpenAI · 2026-04-23 · Frontier · Reasoning
GPT-5.5
OpenAI's response to the V4 / Opus 4.7 surge. Sets the closed-frontier reasoning bar that the open-frontier challengers (V4 Pro, K2.6, GLM-5.1) measure against on SWE-Bench Pro and GPQA. The 5.x family now spans Instant, Thinking, Pro, mini, and nano tiers — useful for matching workload class to model class.
- Model registry ID
- gpt-5-5
Sources openai.com blog, GPT-5.5 system card
Anthropic · 2026-04-07 · Specialist · Agentic
Claude Mythos Preview
The first major model gated by its own developer for safety reasons after independent (UK AISI) red-teaming confirmed autonomous offensive capability. Vendor-risk frameworks now need a capability-gate criterion alongside availability SLAs — diligence on 'what could the next release do that the current one cannot?' is no longer optional.
- Model registry ID
- claude-mythos
Sources anthropic.com, red.anthropic.com, aisi.gov.uk
Anthropic · 2026-04-02 · Frontier · Reasoning
Claude Opus 4.6
Useful as the April 02 baseline against which open-weights eventually moved past mid-month. Procurement teams running 4.5 should plan upgrade to 4.7 directly; 4.6 will likely move to maintenance status before mid-May.
- Model registry ID
- claude-opus-4-6
Sources anthropic.com release notes
Open weights
DeepSeek AI · 2026-04-24 · Open Frontier · Moe
DeepSeek-V4 Pro
The headline release of April. First open-weights model with frontier-class SWE-Bench Pro performance under a permissive license. Procurement teams blocked on closed-source data residency or licensing constraints now have an open frontier alternative for coding workloads. The on-prem floor moved up; the merchant-vs-self-host fork is now real.
- Model registry ID
- deepseek-v4-pro
Sources huggingface.co/deepseek, deepseek.com
DeepSeek AI · 2026-04-24 · Open Frontier · Moe
DeepSeek-V4 Flash
Open-weights answer to GPT-4o-mini and Claude Sonnet at the inference-cost tier. Useful for batch coding agents and high-volume serving where Pro is overkill. Same MIT license, same 1M context, lower active-parameter cost per token.
- Model registry ID
- deepseek-v4-flash
Sources huggingface.co/deepseek model card
Moonshot AI · 2026-04-20 · Open Frontier · Moe
Kimi K2.6
The benchmark moment. K2.6's coding score sits above GPT-5.4 (57.7) and Claude Opus 4.6 (57.3) under an open-weights license. For coding-heavy procurement reads, K2.6 is the new default open-frontier baseline; closed-source premium needs new justification beyond raw score.
- Model registry ID
- kimi-k2-6
Sources moonshot.cn, Artificial Analysis SWE-Bench Pro
Z.AI (Zhipu) · 2026-04-08 · Open Frontier · Agentic
GLM-5.1
Second open-weights model to clear the closed-frontier coding bar this month. The multi-agent-native design point makes GLM-5.1 the open-source reference for agent-of-agents deployments — relevant when coordination overhead matters more than single-shot completion.
- Model registry ID
- glm-5-1
Sources z.ai, GLM-5.1 model card
Alibaba · 2026-04-16 · Open Frontier · Moe
Qwen3.6-35B-A3B
The most deployable of April's open releases on commodity inference fabric. 35B total / 3B active makes Qwen3.6-A3B the open multimodal pick when V4 Pro is too large and K2.6 is single-modality. Pairs naturally with Qwen3.6-Plus on the closed-source side for hybrid deployments.
- Model registry ID
- qwen-3-6-35b-a3b
Sources qwenlm.github.io, model card
DeepSeek AI · 2026-04 · Reasoning · Reasoning
DeepSeek-R2
Same vendor shipping V4 (general) and R2 (reasoning) inside one calendar month is the cadence story. R2 closes the open-vs-closed reasoning gap that R1 opened in 2025 — open-weights reasoning is now a separate procurement category, not a niche.
- Model registry ID
- deepseek-r2
Sources deepseek.com, huggingface.co/deepseek
Architecture watch
Reasoning becomes the default mode, not a separate model.
April's flagships ship reasoning behavior as a default rather than a separate o-series-style fork. Adaptive thinking, where the model decides how much to deliberate, is now the closed-frontier expectation; explicit reasoning toggles look dated by month-end. Procurement consequence: the reasoning-vs-non-reasoning split is collapsing — assume reasoning is on, budget tokens accordingly.
- Examples
- Claude Opus 4.7 (adaptive thinking), GPT-5.5, DeepSeek-R2, GLM-5.1
Sources Vendor model cards (Anthropic, OpenAI, DeepSeek, Z.AI)
Million-token context as baseline.
Every flagship released in April clears the 1M-token threshold. Long-context is no longer a differentiator at the frontier; it is the floor. Workloads that previously required RAG architectures may now be expressible as direct in-context loads — re-evaluate the retrieval layer when re-baselining for V4 / K2.6 / Opus 4.7.
- Examples
- DeepSeek-V4 Pro (1M), Kimi K2.6 (1M+), Claude Opus 4.7 (ultra-long), GPT-5.5 (ultra-long)
Sources Model cards; HuggingFace technical reports
Agentic surface area as a first-class capability.
Agentic isn't a benchmark category anymore; it is a model property declared on the card. Computer-use, MCP support, and multi-agent coordination ship with the flagship rather than as a downstream wrapper. Architectural read: the boundary between 'model' and 'agent runtime' is dissolving — procurement should evaluate the full agent surface, not just the LM weights.
- Examples
- Claude Opus 4.7 (computer-use, MCP, tool-use), GLM-5.1 (multi-agent-native), Qwen3.6-Plus (agentic), Kimi K2.6 (agentic-ready)
Sources Vendor model cards; Model Context Protocol announcements
Capability gating as a market signal.
Anthropic withholding Mythos after UK AISI red-team findings is the first time a major lab gated a frontier-class release on its own initiative for cyber-capability reasons. The signal: capability evaluation now happens before public release, and 'gated' is a status that procurement teams need to track in vendor risk frameworks alongside 'available' and 'deprecated.'
- Examples
- Claude Mythos Preview (withheld)
Sources anthropic.com, red.anthropic.com, aisi.gov.uk
Benchmark moves
SWE-Bench Pro (coding)
Open weights overtook closed for the first time. Two open-weights models clear both closed leaders.
- Kimi K2.6
- 58.6
- GLM-5.1
- 58.4
- GPT-5.4
- 57.7
- Claude Opus 4.6
- 57.3
Sources Artificial Analysis SWE-Bench Pro leaderboard, vendor model cards
Long-context retrieval (1M+ tokens, vendor-reported)
Million-token context lands at the frontier with sub-2% retrieval-error rates across multiple vendors.
- DeepSeek-V4 Pro
- 1M ctx, 1.4% err
- Kimi K2.6
- 1M ctx, 1.7% err
- Claude Opus 4.7
- ultra-long, vendor-reported
- GPT-5.5
- ultra-long, vendor-reported
Sources Vendor technical reports; HuggingFace evaluations
Agentic tool-use (vendor + third-party)
Closed flagships still lead agentic tool-use; open frontier is one tier behind but closing.
- Claude Opus 4.7
- Closed leader
- GPT-5.5
- Closed challenger
- GLM-5.1
- Open leader (multi-agent-native)
- DeepSeek-V4 Pro
- Open challenger
Sources Vendor agent benchmarks; community evaluations
Tier scorecard
As of 2026-04-25
| Tier | Leader | Challenger | Read |
|---|---|---|---|
| Closed frontier | Claude Opus 4.7 | GPT-5.5 | Adaptive thinking is the closed-source differentiator this month. |
| Open frontier | DeepSeek-V4 Pro | Kimi K2.6 | MIT license + 1M context + frontier coding score is the new open baseline. |
| Reasoning | DeepSeek-R2 | GPT-5.5 (thinking mode) | Open-weights reasoning is now its own procurement category. |
| Coding | Kimi K2.6 | GLM-5.1 | Open-weights tops both closed leaders on SWE-Bench Pro this month. |
| Multimodal | Gemini 3.1 Pro | Qwen3.6-Plus | Closed leader from Q1 still holds; Qwen3.6 narrows the gap on the open side. |
| Edge / small | Phi-4-mini | Gemma 4 | Q1 leader carries through April; no major edge-tier release this month. |
Vendor signals
2026-04-21 · Anthropic
Mythos withheld from public access pending capability-gate review.
Sets a precedent for vendor-initiated gating on cyber-capability grounds. Procurement frameworks now need a 'gated' status alongside 'available' / 'deprecated.' Adds a new diligence question: what is the next release the vendor evaluated and chose not to ship?
Sources anthropic.com, red.anthropic.com, aisi.gov.uk
2026-04-24 · DeepSeek AI
Two-config V4 launch (Pro + Flash) plus same-month R2 release under MIT.
Cadence and licensing both shifted. A monthly cycle with permissive licensing across coding, reasoning, and inference-tier variants is the new open-frontier benchmark for vendor velocity. Closed-source vendors with quarterly cadence have a velocity gap, not just a capability gap.
Sources deepseek.com release notes; huggingface.co model cards
2026-04 · OpenAI
GPT-4 family deprecation timeline tightened; 5.x consolidated to five tiers (Instant, Thinking, Pro, mini, nano).
Pricing migration for GPT-4-era workloads accelerates through Q2. Customers should plan migration to 5.x within 60 days; the tier consolidation simplifies model selection but compresses the price-vs-capability slope on the low end.
Sources openai.com platform changelog
2026-04 · Multiple (open-weights serving)
Inference price cuts of 25-40% across major open-weights serving providers following V4 / K2.6 launches.
Open-weights inference economics improved materially this month. Re-baseline cost per token for batch agentic workloads against the new floor; the merchant-vs-self-host break-even moved.
Sources Provider pricing pages; community comparison threads
Watchlist
May 1-7
Closed-frontier response to V4 / K2.6 coding scores.
Anthropic and OpenAI both have credible reasons to push a coding-tuned point release to recover the SWE-Bench Pro lead. Watch for Opus 4.7-Codex or GPT-5.5-Codex within two weeks.
May 1-14
Hyperscaler Q1 prints (MSFT / GOOG / META / AMZN).
AI capex commentary will set the inference-pricing trajectory for the rest of Q2. Tracked in The AI Stack Weekly's capital-flow lens; relevant here as the upstream signal for serving costs.
May 5-20
Mythos status update from Anthropic + AISI.
If gating becomes permanent or extended, 'gated' becomes a durable status class. If gating lifts under conditions, a capability-gate playbook emerges — either way, the procurement read changes.
May 10-25
Open-source reasoning successor cluster.
DeepSeek-R2 set the bar; Qwen, Z.AI, and Moonshot all have R2-class candidates plausibly within 30 days. Watch for the second open-weights reasoning model to clear o-series parity.
May 15-31
First gated-capability standardization proposals.
Mythos creates pressure for a cross-vendor gating taxonomy (capability levels, evaluation protocols, release criteria). Watch for AISI / METR / OpenAI / Anthropic joint statements.
Changelog
- Inaugural issue. Cadence: monthly recap covering April 2026 in one read; weekly cadence begins in May.
- Tree delta sourced from content/llm-tree/models.yaml April 2026 release rows.
- Benchmark scores cite Artificial Analysis leaderboards and vendor model cards as of publish date.
- Scorecard reflects April 2026 close; will refresh weekly as new releases land.