---
title: No new closed frontier shipped; the open/local agent substrate widened underneath it.
publication: The Model Pulse
slug: 2026-W23
issueNumber: 7
isoYear: 2026
isoWeek: 23
cadence: weekly
publishedAt: '2026-06-06'
periodLabel: Week 23 of 2026
canonicalUrl: https://brianletort.ai/industry/models/2026-W23
pdfUrl: https://brianletort.ai/downloads/model-pulse-2026-W23.pdf
schemaVersion: 2026.05.02
treeDelta:
  added:
    - mellum2
    - cosmos-3-nano
    - holo-3-1
  updated: []
frontierMovements:
  - modelId: null
    name: Gemini 3.5 Pro
    vendor: Google DeepMind
    releaseDate: 2026-06 target
    tier: frontier
    architecture: reasoning
openWeights:
  - modelId: mellum2
    name: Mellum2
    vendor: JetBrains
    releaseDate: '2026-06-01'
    tier: edge
    architecture: moe
  - modelId: cosmos-3-nano
    name: NVIDIA Cosmos 3
    vendor: NVIDIA
    releaseDate: '2026-06-01'
    tier: specialist
    architecture: multimodal
  - modelId: holo-3-1
    name: Holo3.1
    vendor: H Company
    releaseDate: '2026-06-02'
    tier: specialist
    architecture: agentic
architecturePatterns:
  - Cheap specialist sub-agents
  - Physical-AI omni-models
  - Frontier release gaps measured in weeks
benchmarks:
  - Artificial Analysis Intelligence Index
  - SWE-Bench Pro
  - Local computer-use deployability
scorecardAsOf: '2026-06-06'
vendorSignals:
  - vendor: Google DeepMind
    date: 2026-06
    signal: >-
      Gemini 3.5 Pro remains promised for June with no public API model ID, pricing, or third-party
      benchmark row by W23 close
  - vendor: JetBrains
    date: '2026-06-01'
    signal: Released Mellum2 under Apache 2.0 with an explicit low-latency production-workload positioning
  - vendor: NVIDIA
    date: '2026-06-01'
    signal: Released Cosmos 3 on Hugging Face while also announcing Vera Rubin production at GTC Taipei
---

# No new closed frontier shipped; the open/local agent substrate widened underneath it.

*The Model Pulse · Issue 07 · Week 23 of 2026 · Published 2026-06-06*

## Big Read

W23 did not produce the expected Gemini 3.5 Pro GA or a fresh Anthropic/OpenAI frontier release. That absence is the story: Claude Opus 4.8 remains the public closed-frontier leader for coding and agentic work, while Google kept Pro in the June watch window and the model layer's actual shipping activity moved down-stack. JetBrains released Mellum2, an Apache-2.0 12B/2.5B-active MoE designed for low-latency routing, RAG, summarization, validation, and sub-agent calls; NVIDIA released Cosmos 3 as an open physical-AI omni-model with Nano 16B and Super 64B variants; H Company released Holo3.1 with local computer-use sizes and quantized checkpoints. The procurement implication is sharper than another leaderboard reshuffle: production agent systems are becoming portfolios of models. Keep Opus/GPT/Gemini-class models for high-risk reasoning and codebase-scale orchestration, but push cheap, private, repeated sub-agent work into specialized open/local models. The tree delta therefore adds efficient-agent and physical-AI nodes rather than another general chatbot crown.

## Tree delta

Three W23 additions: Mellum2 for efficient text/code sub-agent workloads, Cosmos 3 for physical-AI omni-modeling, and Holo3.1 for local computer-use agents.

**Added (3):** `mellum2`, `cosmos-3-nano`, `holo-3-1`.

**Updated:** none.

*Gemini 3.5 Pro remains pending; no tree row is added until Google publishes the GA model card or API identifier.*

## Frontier movements

### Gemini 3.5 Pro

**Google DeepMind** · 2026-06 target · frontier · reasoning

_Still pending at W23 close: Google has said Pro follows Gemini 3.5 Flash in June, but no public API ID, pricing, or independent benchmark row landed in-window_

This is a frontier movement by absence. Buyers waiting for Gemini 3.5 Pro should keep the launch on the June watchlist but should not pause current coding-agent baselines: the public board still has Opus 4.8 leading the closed frontier, and Pro's economics and benchmark profile remain unverified. The likely enterprise split is task routing — Gemini for huge-context/multimodal work, Opus/GPT for coding and agentic reliability — not a universal replacement.

Source: [Google Gemini 3.5 announcement; AI Tool Bolt June comparison](https://blog.google/intl/en-africa/products/explore-get-answers/gemini-3-5/).

## Open weights

### Mellum2 `mellum2`

**JetBrains** · 2026-06-01 · edge · moe

_Apache-2.0 12B MoE with 2.5B active parameters per token for low-latency text/code sub-agent workloads_

Mellum2 is not trying to win the frontier leaderboard; it is trying to lower the cost of the thousands of routine model calls inside agent systems. Routing, RAG, summarization, validation, and lightweight code tasks are exactly where private deployments want an efficient open model. If the benchmark claims hold, this is a practical procurement node for teams trying to cut orchestration cost without sending every step to a closed flagship.

Source: [Hugging Face JetBrains Mellum2 launch](https://huggingface.co/blog/JetBrains/mellum2-launch).

### NVIDIA Cosmos 3 `cosmos-3-nano`

**NVIDIA** · 2026-06-01 · specialist · multimodal

_Open physical-AI omni-model family combining world generation, physical reasoning, and action generation in Nano 16B and Super 64B variants_

Cosmos 3 widens the model tree away from text-only agents and into physical AI. The important shift is a unified model that can reason over world state and action generation rather than stitching separate generation and control pipelines together. Robotics, simulation, and industrial automation teams should evaluate it as synthetic-data and reasoning infrastructure, not as a chatbot substitute.

Source: [Hugging Face NVIDIA Cosmos 3 launch](https://huggingface.co/blog/nvidia/cosmos-3-for-physical-ai).

### Holo3.1 `holo-3-1`

**H Company** · 2026-06-02 · specialist · agentic

_Local computer-use agent family with 0.8B, 4B, 9B, and 35B-A3B sizes plus FP8, Q4 GGUF, and NVFP4 checkpoints_

Holo3.1 pushes computer-use agents toward deployability: multiple sizes, quantized checkpoints, and local inference targets matter more than a single headline score. Enterprises automating browser, desktop, and internal-tool workflows can now separate privacy-sensitive UI action from hosted frontier reasoning. That supports a two-layer architecture: local CUA for execution, closed frontier for planning and verification.

Source: [Hugging Face Holo3.1 launch](https://huggingface.co/blog/hcompany/holo31).

## Architecture watch

### Cheap specialist sub-agents

_Examples:_ Mellum2, Holo3.1-0.8B / 4B / 9B, Claude Opus 4.8 fast mode.

The agent stack is splitting into high-reasoning planners and cheap repeated workers. Mellum2 and Holo3.1 are purpose-built for the calls that happen hundreds or thousands of times inside a workflow: routing, validation, summarization, UI action, and local execution. Model routers should now budget by step type rather than treating one flagship as the default for every agent call.

Source: Hugging Face Mellum2 and Holo3.1 launches.

### Physical-AI omni-models

_Examples:_ NVIDIA Cosmos 3 Nano, NVIDIA Cosmos 3 Super, Bernini-R renderer.

Open model activity is expanding from language and code into physical-world simulation, video rendering, and action generation. Cosmos 3's combined world generation, physical reasoning, and action generation points toward a branch where synthetic data and robotics workflows become first-class model workloads. That is a different buyer and deployment path than enterprise chat.

Source: Hugging Face NVIDIA Cosmos 3 launch; ByteDance Bernini-R Hugging Face card.

### Frontier release gaps measured in weeks

_Examples:_ Claude Opus 4.8, Gemini 3.5 Pro pending, GPT-5.6 speculation.

The absence of a new frontier release this week matters because expectations have compressed. Buyers are now tempted to delay procurement for a model that may arrive in days. The practical answer is to separate infrastructure choices from model choice: standardize evaluation harnesses, routers, and cost controls so a June GA can be tested and slotted without freezing current deployments.

Source: Google Gemini 3.5 announcement; Anthropic Opus 4.8 announcement.

## Benchmark moves

### Artificial Analysis Intelligence Index

No new W23 leaderboard reset; Claude Opus 4.8 remains the public #1 at 61.4 while Gemini 3.5 Pro is still pending GA

  - Claude Opus 4.8: 61.4
  - GPT-5.5: ~60
  - Gemini 3.1 Pro: ~57

Source: [Artificial Analysis summaries via June model comparisons](https://aitoolbolt.com/claude-opus-4-8-vs-gemini-3-5-pro/).

### SWE-Bench Pro

Closed frontier still leads coding: Opus 4.8's 69.2% remains the public bar; no new open release in W23 changes the top coding score

  - Claude Opus 4.8: 69.2%
  - GPT-5.5: ~66-67%
  - Gemini 3.1 Pro: ~62%

Source: [AI Tool Bolt / Pristren June benchmark summaries](https://pristren.com/blog/claude-opus-4-8-vs-gpt-5-5-gemini-3-1-june-2026/).

### Local computer-use deployability

Holo3.1 shifts the measurable axis from one top-line CUA score to model size and quantization availability for local execution

  - Holo3.1-0.8B: ultra-light local
  - Holo3.1-9B: balanced local
  - Holo3.1-35B-A3B: state-of-the-art tier

Source: [Hugging Face Holo3.1 launch](https://huggingface.co/blog/hcompany/holo31).

## Tier scorecard

_As of 2026-06-06._

| Tier | Leader | Challenger | Note |
|---|---|---|---|
| Closed frontier | Claude Opus 4.8 | GPT-5.5 | No W23 reset; Gemini 3.5 Pro remains the watched June challenger rather than a published benchmark row. |
| Open frontier | DeepSeek V4-Pro | GLM-5.1 | No new open frontier text model displaced the April leaders; W23 open activity shifted to specialist/local models. |
| Reasoning | Claude Opus 4.8 | GPT-5.5 | Closed reasoning leadership steady while Gemini 3.5 Pro remains pending. |
| Coding | Claude Opus 4.8 | GPT-5.5 | Opus 4.8 still owns the visible SWE-Bench Pro lead; Mellum2 matters for cheap sub-agent code/text calls. |
| Multimodal | Gemini 3.1 Pro | Cosmos 3 | Gemini remains the general multimodal reference; Cosmos 3 creates a specialist physical-AI branch. |
| Edge / small | Mellum2 | Holo3.1-9B | Efficient local/sub-agent models were the week's real release activity. |

## Vendor signals

- **2026-06 · Google DeepMind — Gemini 3.5 Pro remains promised for June with no public API model ID, pricing, or third-party benchmark row by W23 close**
  - Procurement teams should prepare an eval slot but avoid freezing current deployments on an unreleased model. The operating pattern is rapid re-baselining, not launch-date speculation.
  - _Source:_ [Google Gemini 3.5 announcement and June developer guidance](https://blog.google/intl/en-africa/products/explore-get-answers/gemini-3-5/)
- **2026-06-01 · JetBrains — Released Mellum2 under Apache 2.0 with an explicit low-latency production-workload positioning**
  - IDE and enterprise-platform vendors can now point to a plausible open/private model for background text-code tasks. That increases pressure on hosted copilots to justify every closed-frontier call by risk or quality, not habit.
  - _Source:_ [Hugging Face JetBrains Mellum2 launch](https://huggingface.co/blog/JetBrains/mellum2-launch)
- **2026-06-01 · NVIDIA — Released Cosmos 3 on Hugging Face while also announcing Vera Rubin production at GTC Taipei**
  - NVIDIA is binding the model and infrastructure stories together: physical-AI models create demand for the simulation, synthetic-data, and rack-scale compute stack it sells. Buyers should evaluate model capability and deployment substrate together.
  - _Source:_ [Hugging Face Cosmos 3 launch; NVIDIA GTC Taipei](https://huggingface.co/blog/nvidia/cosmos-3-for-physical-ai)

## Watchlist

- **Jun 7-30 — Gemini 3.5 Pro GA.** The first public model card, price row, API ID, and Artificial Analysis pass will determine whether June becomes a true frontier reset or just a routing expansion.
- **Jun-Jul — Mythos-class Anthropic availability.** Anthropic has publicly framed stronger gated models as pending cyber safeguards. A wider release would change the closed-frontier scorecard more than another Opus point release.
- **Jun-Aug — Open/local agent model adoption.** Downloads, integrations, and benchmark replications for Mellum2, Cosmos 3, and Holo3.1 will show whether specialist open models are becoming production substrate or just launch-week noise.

## Changelog

- Added Mellum2, Cosmos 3 Nano, and Holo3.1 to the LLM tree and reframed W23 around efficient/local specialist models rather than a closed-frontier release.

---

Source of truth: `src/data/industry/models/2026-W23.ts`. Canonical HTML: <https://brianletort.ai/industry/models/2026-W23>. PDF: <https://brianletort.ai/downloads/model-pulse-2026-W23.pdf>. Tree: <https://brianletort.ai/industry/tree>.
