Skip to content

Model layer

The Model Pulse

For architects tracking model capability shifts.

No new closed frontier shipped; the open/local agent substrate widened underneath it.

Big read

W23 did not produce the expected Gemini 3.5 Pro GA or a fresh Anthropic/OpenAI frontier release. That absence is the story: Claude Opus 4.8 remains the public closed-frontier leader for coding and agentic work, while Google kept Pro in the June watch window and the model layer's actual shipping activity moved down-stack. JetBrains released Mellum2, an Apache-2.0 12B/2.5B-active MoE designed for low-latency routing, RAG, summarization, validation, and sub-agent calls; NVIDIA released Cosmos 3 as an open physical-AI omni-model with Nano 16B and Super 64B variants; H Company released Holo3.1 with local computer-use sizes and quantized checkpoints. The procurement implication is sharper than another leaderboard reshuffle: production agent systems are becoming portfolios of models. Keep Opus/GPT/Gemini-class models for high-risk reasoning and codebase-scale orchestration, but push cheap, private, repeated sub-agent work into specialized open/local models. The tree delta therefore adds efficient-agent and physical-AI nodes rather than another general chatbot crown.

Tree delta

Three W23 additions: Mellum2 for efficient text/code sub-agent workloads, Cosmos 3 for physical-AI omni-modeling, and Holo3.1 for local computer-use agents.

Registry movement

Gemini 3.5 Pro remains pending; no tree row is added until Google publishes the GA model card or API identifier.

Added
mellum2, cosmos-3-nano, holo-3-1
Updated
None

Frontier movements

Google DeepMind · 2026-06 target · Frontier · Reasoning

Gemini 3.5 Pro

This is a frontier movement by absence. Buyers waiting for Gemini 3.5 Pro should keep the launch on the June watchlist but should not pause current coding-agent baselines: the public board still has Opus 4.8 leading the closed frontier, and Pro's economics and benchmark profile remain unverified. The likely enterprise split is task routing — Gemini for huge-context/multimodal work, Opus/GPT for coding and agentic reliability — not a universal replacement.

Sources Google Gemini 3.5 announcement; AI Tool Bolt June comparison

Open weights

JetBrains · 2026-06-01 · Edge · Moe

Mellum2

Mellum2 is not trying to win the frontier leaderboard; it is trying to lower the cost of the thousands of routine model calls inside agent systems. Routing, RAG, summarization, validation, and lightweight code tasks are exactly where private deployments want an efficient open model. If the benchmark claims hold, this is a practical procurement node for teams trying to cut orchestration cost without sending every step to a closed flagship.

Model registry ID
mellum2

Sources Hugging Face JetBrains Mellum2 launch

NVIDIA · 2026-06-01 · Specialist · Multimodal

NVIDIA Cosmos 3

Cosmos 3 widens the model tree away from text-only agents and into physical AI. The important shift is a unified model that can reason over world state and action generation rather than stitching separate generation and control pipelines together. Robotics, simulation, and industrial automation teams should evaluate it as synthetic-data and reasoning infrastructure, not as a chatbot substitute.

Model registry ID
cosmos-3-nano

Sources Hugging Face NVIDIA Cosmos 3 launch

H Company · 2026-06-02 · Specialist · Agentic

Holo3.1

Holo3.1 pushes computer-use agents toward deployability: multiple sizes, quantized checkpoints, and local inference targets matter more than a single headline score. Enterprises automating browser, desktop, and internal-tool workflows can now separate privacy-sensitive UI action from hosted frontier reasoning. That supports a two-layer architecture: local CUA for execution, closed frontier for planning and verification.

Model registry ID
holo-3-1

Sources Hugging Face Holo3.1 launch

Architecture watch

Cheap specialist sub-agents

The agent stack is splitting into high-reasoning planners and cheap repeated workers. Mellum2 and Holo3.1 are purpose-built for the calls that happen hundreds or thousands of times inside a workflow: routing, validation, summarization, UI action, and local execution. Model routers should now budget by step type rather than treating one flagship as the default for every agent call.

Examples
Mellum2, Holo3.1-0.8B / 4B / 9B, Claude Opus 4.8 fast mode

Sources Hugging Face Mellum2 and Holo3.1 launches

Physical-AI omni-models

Open model activity is expanding from language and code into physical-world simulation, video rendering, and action generation. Cosmos 3's combined world generation, physical reasoning, and action generation points toward a branch where synthetic data and robotics workflows become first-class model workloads. That is a different buyer and deployment path than enterprise chat.

Examples
NVIDIA Cosmos 3 Nano, NVIDIA Cosmos 3 Super, Bernini-R renderer

Sources Hugging Face NVIDIA Cosmos 3 launch; ByteDance Bernini-R Hugging Face card

Frontier release gaps measured in weeks

The absence of a new frontier release this week matters because expectations have compressed. Buyers are now tempted to delay procurement for a model that may arrive in days. The practical answer is to separate infrastructure choices from model choice: standardize evaluation harnesses, routers, and cost controls so a June GA can be tested and slotted without freezing current deployments.

Examples
Claude Opus 4.8, Gemini 3.5 Pro pending, GPT-5.6 speculation

Sources Google Gemini 3.5 announcement; Anthropic Opus 4.8 announcement

Benchmark moves

SWE-Bench Pro

Closed frontier still leads coding: Opus 4.8's 69.2% remains the public bar; no new open release in W23 changes the top coding score

Claude Opus 4.8
69.2%
GPT-5.5
~66-67%
Gemini 3.1 Pro
~62%

Sources AI Tool Bolt / Pristren June benchmark summaries

Local computer-use deployability

Holo3.1 shifts the measurable axis from one top-line CUA score to model size and quantization availability for local execution

Holo3.1-0.8B
ultra-light local
Holo3.1-9B
balanced local
Holo3.1-35B-A3B
state-of-the-art tier

Sources Hugging Face Holo3.1 launch

Tier scorecard

As of 2026-06-06

TierLeaderChallengerRead
Closed frontierClaude Opus 4.8GPT-5.5No W23 reset; Gemini 3.5 Pro remains the watched June challenger rather than a published benchmark row.
Open frontierDeepSeek V4-ProGLM-5.1No new open frontier text model displaced the April leaders; W23 open activity shifted to specialist/local models.
ReasoningClaude Opus 4.8GPT-5.5Closed reasoning leadership steady while Gemini 3.5 Pro remains pending.
CodingClaude Opus 4.8GPT-5.5Opus 4.8 still owns the visible SWE-Bench Pro lead; Mellum2 matters for cheap sub-agent code/text calls.
MultimodalGemini 3.1 ProCosmos 3Gemini remains the general multimodal reference; Cosmos 3 creates a specialist physical-AI branch.
Edge / smallMellum2Holo3.1-9BEfficient local/sub-agent models were the week's real release activity.

Vendor signals

2026-06 · Google DeepMind

Gemini 3.5 Pro remains promised for June with no public API model ID, pricing, or third-party benchmark row by W23 close

Procurement teams should prepare an eval slot but avoid freezing current deployments on an unreleased model. The operating pattern is rapid re-baselining, not launch-date speculation.

Sources Google Gemini 3.5 announcement and June developer guidance

2026-06-01 · JetBrains

Released Mellum2 under Apache 2.0 with an explicit low-latency production-workload positioning

IDE and enterprise-platform vendors can now point to a plausible open/private model for background text-code tasks. That increases pressure on hosted copilots to justify every closed-frontier call by risk or quality, not habit.

Sources Hugging Face JetBrains Mellum2 launch

2026-06-01 · NVIDIA

Released Cosmos 3 on Hugging Face while also announcing Vera Rubin production at GTC Taipei

NVIDIA is binding the model and infrastructure stories together: physical-AI models create demand for the simulation, synthetic-data, and rack-scale compute stack it sells. Buyers should evaluate model capability and deployment substrate together.

Sources Hugging Face Cosmos 3 launch; NVIDIA GTC Taipei

Watchlist

Jun 7-30

Gemini 3.5 Pro GA

The first public model card, price row, API ID, and Artificial Analysis pass will determine whether June becomes a true frontier reset or just a routing expansion.

Jun-Jul

Mythos-class Anthropic availability

Anthropic has publicly framed stronger gated models as pending cyber safeguards. A wider release would change the closed-frontier scorecard more than another Opus point release.

Jun-Aug

Open/local agent model adoption

Downloads, integrations, and benchmark replications for Mellum2, Cosmos 3, and Holo3.1 will show whether specialist open models are becoming production substrate or just launch-week noise.

Changelog

  • Added Mellum2, Cosmos 3 Nano, and Holo3.1 to the LLM tree and reframed W23 around efficient/local specialist models rather than a closed-frontier release.