{
  "_meta": {
    "publication": "The Model Pulse",
    "schemaVersion": "2026.05.02",
    "generatedAt": "2026-09-08T22:45:10.939Z",
    "canonicalUrl": "https://brianletort.ai/industry/models/2026-W20",
    "markdownUrl": "https://brianletort.ai/industry/models/2026-W20/llm.md",
    "pdfUrl": "https://brianletort.ai/downloads/model-pulse-2026-W20.pdf",
    "treeUrl": "https://brianletort.ai/industry/tree",
    "sourceFile": "src/data/industry/models/2026-W20.ts"
  },
  "issue": {
    "slug": "2026-W20",
    "isoYear": 2026,
    "isoWeek": 20,
    "issueNumber": 4,
    "publishedAt": "2026-05-17",
    "cadence": "weekly",
    "periodLabel": "Week 20 of 2026",
    "bigRead": {
      "headline": "Frontier text took a breath; the specialist canopy widened — video, on-device multimodal, open world-models — and agent-runtime monetization became the story.",
      "body": "Frontier text models took a breath this week — no new GPT-class or Claude refresh between W19 and Google I/O (May 19-20), and LMArena top-of-board moved less than 1 Elo. But the specialist canopy widened sharply, with three model rows joining the tree across video / embodied, on-device multimodal, and open-source world modeling.\n\nPerceptron Mk1 collapsed video and embodied reasoning costs by 80-90% on May 12, pricing frontier-grade physical-AI perception under Gemini Flash Lite ($0.15 / $1.50 per Mtok, 85.1 EmbSpatialBench, 72.4 RefSpatialBench) and forcing a re-cost of any 2H roadmap that ships robotics, surveillance, or screen-watching agents. NVIDIA's open-weights SANA-WM (May 15, Apache 2.0) put a minute-scale 720p world model on a single RTX 5090 in 34 seconds — robotics, AV simulation, and synthetic-data teams can now in-house what they previously rented from closed video APIs. OpenBMB's MiniCPM-V 4.6 1.3B (May 11) reset the low end of multimodal SLMs at 262K context with native iOS / Android / HarmonyOS deployment, compressing on-device feature timelines from quarters to weeks.\n\nOn the platform layer, the unit of competition shifted from model intelligence to agent-runtime monetization. Anthropic's Code with Claude conference (May 6-12) made Managed Agents — Dreaming (between-session memory consolidation), Multiagent Orchestration, Outcomes (rubric-graded), and signed Webhooks — a hosted service; Claude Platform on AWS hit GA May 11; xAI shipped Grok Build CLI on May 14 with a SuperGrok Heavy tier at $300/mo; OpenAI's realtime voice trio (GPT-Realtime-2 / Translate / Whisper, added W19) went broadly available May 11 with 128K context and adjustable reasoning effort. Architects building bespoke agent stacks should justify build vs buy explicitly in 2H plans.\n\nThe W18 Pentagon vendor-disqualification thread compounded. UK AISI published paired evaluations May 13 confirming autonomous offensive-cyber capability is now doubling every 4.7 months (aligned with METR's 4.2-month SWE figure), and Anthropic emailed Max-20x subscribers a June 15 policy that moves Claude Code third-party agents off subscription rate limits — 12x-175x effective price increase per workload. Capability gating, federal procurement, and pricing structure are now durable vendor-risk dimensions; bench scores alone no longer settle a procurement decision.\n\nFor the May 19-26 window, the load-bearing catalysts are Google I/O 2026 (Gemini 3.2 Flash / Pro confirmation, fabric session), NVIDIA Q1 FY27 (May 20), Microsoft Build (May 19-22), the Samsung HBM4 walkout starting May 21, and any Anthropic / OpenAI counter-launch following I/O."
    },
    "treeDelta": {
      "summary": "3 model rows added across 3 vendors in a single 7-day window; the canopy widened in specialist multimodal (video / embodied), open-source world modeling, and on-device multimodal SLMs. Closed-source text frontier was quiet ahead of Google I/O.",
      "added": [
        "minicpm-v-4-6",
        "perceptron-mk1",
        "sana-wm"
      ],
      "updated": [
        "gpt-realtime-2",
        "gpt-realtime-translate",
        "gpt-realtime-whisper"
      ],
      "note": "OpenAI realtime voice family (added W19) hit broad public availability May 11 — flipped status to GA in updated rather than re-adding. Frontier text held flat: GPT-5.5 (xhigh) at AA Index 60 unchanged; Claude Opus 4.7 (Adaptive Max) at 57; LMArena top-of-board moved less than 1 Elo."
    },
    "frontierMovements": [
      {
        "modelId": "perceptron-mk1",
        "name": "Perceptron Mk1",
        "vendor": "Perceptron",
        "releaseDate": "2026-05-12",
        "headline": "Re-cost any 2H roadmap that ships video, robotics, or screen-watching agents — frontier-grade perception just landed below Gemini Flash Lite.",
        "why": "Perceptron Mk1 matches Gemini Pro / Claude / GPT-class on video and embodied reasoning at $0.15 / $1.50 per Mtok — 80-90% under incumbent multimodal pricing — and scores 85.1 EmbSpatialBench / 72.4 RefSpatialBench. Architects building real-time physical-AI pipelines (manufacturing QA, surveillance, robotics policy verification) should pilot it this quarter; procurement teams should treat it as a credible second source against Google for high-volume video workloads.",
        "tier": "specialist",
        "architecture": "multimodal",
        "source": "perceptron.inc/blog/introducing-perceptron-mk1, VentureBeat, BusinessWire"
      },
      {
        "modelId": "gpt-realtime-2",
        "name": "GPT-Realtime-2 / Translate / Whisper (suite GA)",
        "vendor": "OpenAI",
        "releaseDate": "2026-05-11",
        "headline": "Voice agents now get GPT-5-class reasoning at production scale — refresh any call-center or IVR roadmap that still assumes pre-reasoning realtime APIs.",
        "why": "OpenAI's realtime trio shifted from W19 announcement to broad API availability on May 11 with 128K context (up from 32K), adjustable reasoning effort, and parallel tool calls with audible status. Procurement should re-RFP call-center and accessibility workflows that were locked in at last-gen realtime pricing; competitors (Anthropic, Google, Inworld) will have to answer on reasoning-in-voice within 90 days.",
        "tier": "frontier",
        "architecture": "multimodal",
        "source": "openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api"
      }
    ],
    "openWeights": [
      {
        "modelId": "minicpm-v-4-6",
        "name": "MiniCPM-V 4.6 1.3B",
        "vendor": "OpenBMB",
        "releaseDate": "2026-05-11",
        "headline": "On-device multimodal just got serious — push a tier of vision / OCR / screen-understanding workloads off cloud APIs before Q3 budget locks.",
        "why": "1.3B params, 262K context, runs natively on iOS / Android / HarmonyOS with quantized variants (GGUF, BNB, AWQ, GPTQ) shipped day one. Hits AA Intelligence Index 13 — beats Qwen3.5-0.8B at 19x lower token cost and matches Qwen3.5 2B on many vision tasks. Operators should pilot offline vision agents (warehouse, retail, field service) and reconsider any cloud-locked on-screen-understanding contract; investors should mark another point on the SLM / edge cost curve.",
        "tier": "edge",
        "architecture": "multimodal",
        "source": "artificialanalysis.ai/articles/openbmb-launches-minicpm-v-4-6-1-3b-instruct, huggingface.co/openbmb/MiniCPM-V-4.6"
      },
      {
        "modelId": "sana-wm",
        "name": "SANA-WM",
        "vendor": "NVIDIA",
        "releaseDate": "2026-05-15",
        "headline": "Open-weight world models now fit on one GPU — robotics, simulation, and synthetic-data teams can in-house what they previously rented from closed video APIs.",
        "why": "2.6B Hybrid Linear Diffusion Transformer that produces 60-second 720p clips with 6-DoF camera control in 34 seconds on a single RTX 5090 (NVFP4), a 36x throughput gain over prior open baselines. Architects building robotics pretraining, AV simulation, or content-pipeline tooling should fork SANA-WM rather than sign multi-year deals with closed video APIs; procurement should ask their current video-gen vendor what justifies a 36x throughput premium.",
        "tier": "specialist",
        "architecture": "multimodal",
        "source": "arxiv.org/abs/2605.15178, NVlabs/Sana GitHub"
      }
    ],
    "architectureWatch": [
      {
        "pattern": "Specialist video / embodied tier breaks out at frontier quality, sub-Flash-Lite cost",
        "examples": [
          "Perceptron Mk1 (closed, May 12)",
          "MolmoAct 2 (Ai2, May 5 — adjacent)"
        ],
        "body": "Two physical-AI releases inside the 14-day grace window (Perceptron Mk1 closed, MolmoAct 2 open) signal that video / embodied reasoning is splitting off from general multimodal frontier and pricing under it. Procurement teams that lumped these workloads into a Gemini or Claude SKU should re-segment them in 2H planning; architects should expect a two-tier multimodal stack (general LMM + specialist physical AI) by year-end.",
        "source": "venturebeat.com/technology/perceptron-mk1-shocks-with-highly-performant-video-analysis-ai-model"
      },
      {
        "pattern": "Hybrid linear attention compresses long-context cost without quality cliff",
        "examples": [
          "SANA-WM (Hybrid Linear Diffusion Transformer)",
          "Subquadratic SubQ (May 5, 12M context, verification pending)"
        ],
        "body": "NVIDIA's SANA-WM uses frame-wise Gated DeltaNet plus softmax attention to deliver minute-scale video on one GPU; Subquadratic's SubQ pushes the same idea to 12M-token text context with claimed ~1000x compute reduction. Operators with long-document, long-video, or long-session agentic workloads should add at least one subquadratic-attention model to their 2H pilot list — the cost curve for ultra-long context is moving faster than transformer-only roadmaps assume.",
        "source": "arxiv.org/abs/2605.15178; subquadratic.ai"
      },
      {
        "pattern": "Edge multimodal SLMs ship native quantization + OS coverage out of the box",
        "examples": [
          "MiniCPM-V 4.6 1.3B"
        ],
        "body": "OpenBMB shipped GGUF / BNB / AWQ / GPTQ quantizations and reference deployments for iOS, Android, and HarmonyOS on day one — the friction between 'open-weights release' and 'phone build' has collapsed. Operators planning on-device features can compress timelines from quarters to weeks; vendors who still ship weights without quantization or mobile runtime support look behind the curve.",
        "source": "huggingface.co/openbmb/MiniCPM-V-4.6"
      },
      {
        "pattern": "Managed agent runtimes (memory, orchestration, outcome graders) move from preview to hosted product",
        "examples": [
          "Anthropic Claude Managed Agents — Dreaming, Multiagent Orchestration, Outcomes, Webhooks"
        ],
        "body": "Anthropic's Code with Claude (May 6-12) made Dreaming (between-session memory curation), coordinator-subagent orchestration, and rubric-graded Outcomes a hosted service — and Claude Platform on AWS hit GA May 11. Architects building agent stacks should stop rolling these primitives in-house unless they have a clear differentiator; procurement should ask their LLM vendor for an explicit answer on agent-runtime feature parity.",
        "source": "claude.com/blog/code-w-claude-sf-2026-sf"
      }
    ],
    "benchmarkMoves": [
      {
        "benchmark": "SWE-Bench Pro (agentic coding)",
        "movement": "Claude Mythos Preview extends its lead with Anthropic occupying the top two slots; GPT-5.5 trails by ~19 points — re-anchor coding-agent procurement on Claude unless GPT-5.5 Instant variant changes pricing.",
        "rows": [
          {
            "model": "Claude Mythos Preview (Anthropic, gated)",
            "score": "77.8"
          },
          {
            "model": "Claude Opus 4.7 (Adaptive)",
            "score": "64.3"
          },
          {
            "model": "GPT-5.5",
            "score": "58.6"
          },
          {
            "model": "Kimi K2.6 / GLM-5.1 / MiMo V2.5 Pro",
            "score": "57-59"
          }
        ],
        "source": "benchlm.ai/benchmarks/swePro"
      },
      {
        "benchmark": "Artificial Analysis Intelligence Index v4 (composite)",
        "movement": "GPT-5.5 xhigh leads at 60; GPT-5.5 high 59; Claude Opus 4.7 (Adaptive Max) 57 — Anthropic and OpenAI within a 3-point band, no clear frontier text leader heading into Google I/O.",
        "rows": [
          {
            "model": "GPT-5.5 (xhigh)",
            "score": "60"
          },
          {
            "model": "GPT-5.5 (high)",
            "score": "59"
          },
          {
            "model": "Claude Opus 4.7 (Adaptive Max)",
            "score": "57"
          },
          {
            "model": "Kimi K2.6 / MiMo V2.5 Pro (open top)",
            "score": "54"
          }
        ],
        "source": "artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index"
      },
      {
        "benchmark": "LMArena overall text (May-16 snapshot)",
        "movement": "Anthropic occupies top 5 (Opus 4.6 Thinking, Opus 4.6, Opus 4.7, Opus 4.7 Thinking, Sonnet 4.6); GPT-5.5 at #6 / #7 — text leaderboard moved less than 1 Elo this week, confirming a quiet text-frontier holding pattern.",
        "rows": [
          {
            "model": "Claude Opus 4.6 Thinking",
            "score": "1525"
          },
          {
            "model": "Claude Opus 4.7",
            "score": "1515"
          },
          {
            "model": "GPT-5.5",
            "score": "1490"
          },
          {
            "model": "Kimi K2.6",
            "score": "1457"
          }
        ],
        "source": "huggingface.co/datasets/lmarena-ai/leaderboard-dataset"
      },
      {
        "benchmark": "EmbSpatialBench / RefSpatialBench (embodied spatial reasoning)",
        "movement": "Perceptron Mk1 sets a new ceiling for affordable spatial reasoning — leaders on closed multimodal bench at a fraction of incumbent pricing, opening physical-AI workloads that previously needed Gemini Pro / Claude Opus.",
        "rows": [
          {
            "model": "Perceptron Mk1 (EmbSpatialBench)",
            "score": "85.1"
          },
          {
            "model": "Perceptron Mk1 (RefSpatialBench)",
            "score": "72.4"
          }
        ],
        "source": "venturebeat.com/technology/perceptron-mk1-shocks-with-highly-performant-video-analysis-ai-model"
      }
    ],
    "scorecard": {
      "asOf": "2026-05-17",
      "rows": [
        {
          "tier": "Closed frontier",
          "leader": "GPT-5.5 (xhigh) — AA Index 60",
          "challenger": "Claude Opus 4.7 (Adaptive Max) — AA Index 57; LMArena leader",
          "note": "Three-point band at the top — no clear winner ahead of Google I/O; default by use case, not vendor."
        },
        {
          "tier": "Open frontier",
          "leader": "GLM-5 (Z.ai, 744B / 40B active) — top open-source on Vending Bench 2",
          "challenger": "DeepSeek V4 Pro (1.6T / 49B active) — V4.1 with full-modal coverage scheduled June",
          "note": "Open frontier flat this week; DeepSeek V4.1 in June is the next catalyst for the open-vs-closed lever."
        },
        {
          "tier": "Reasoning",
          "leader": "Claude Opus 4.7 — leads GDPval-AA at 1753 Elo",
          "challenger": "Kimi K2.6 — top open-weight reasoning at 1457 LMArena Elo",
          "note": "Anthropic still owns enterprise reasoning by ~80+ Elo; Moonshot leads the open side."
        },
        {
          "tier": "Coding",
          "leader": "Claude Mythos Preview — 77.8 SWE-Bench Pro (BenchLM)",
          "challenger": "Claude Opus 4.7 (Adaptive) — 64.3 SWE-Bench Pro; GPT-5.5 at 58.6",
          "note": "Anthropic owns the top two coding slots; Mistral Medium 3.5 (77.6 SWE-Verified) leads the open challenger track."
        },
        {
          "tier": "Multimodal",
          "leader": "Gemini 3.1 Pro — 57.2 AA Index, top vision arena",
          "challenger": "Perceptron Mk1 (video / embodied specialist) — Gemini-Pro-class at 80-90% lower cost",
          "note": "General LMM leader unchanged; specialist multimodal split is forming under it (video / embodied + on-device)."
        },
        {
          "tier": "Edge / small",
          "leader": "Gemma 4 E4B / 31B Dense — HF top trending, #3 LMArena open-leader",
          "challenger": "MiniCPM-V 4.6 1.3B — on-device multimodal with 262K context, native iOS / Android",
          "note": "Edge tier added a credible multimodal SLM this week; sub-2B with 200K+ context is the new normal."
        }
      ]
    },
    "vendorSignals": [
      {
        "vendor": "Anthropic",
        "date": "2026-05-13",
        "signal": "Max-20x policy change effective June 15: Claude Agent SDK, `claude -p`, GitHub Actions, third-party agents move to separate $20-$200/mo metered credit at API list prices — 12x-175x effective price increase per workload",
        "meaning": "Ends Claude Code subscription arbitrage. Architects running Claude Code at scale must model true API-rate burn before the June 15 cutover or shift workloads to alternative coding agents (OpenAI Codex, xAI Grok Build, Cursor). Procurement should reopen agent-runtime contracts assuming Anthropic's monetization stance is structural; the third-party agent ecosystem (OpenClaw, T3 Code, Conductor, Zed, Jean) faces a churn test.",
        "source": "support.claude.com, VentureBeat, XDA-Developers"
      },
      {
        "vendor": "Anthropic",
        "date": "2026-05-11",
        "signal": "Claude Platform on AWS hits GA",
        "meaning": "Enterprise procurement teams already standardized on AWS can now buy Claude (including Managed Agents) through AWS billing, IAM, and authentication — removes the last friction excuse not to deploy Claude alongside or instead of Bedrock-native models; expect Anthropic enterprise pipeline to compress this quarter.",
        "source": "claude.com/blog/claude-platform-on-aws"
      },
      {
        "vendor": "Anthropic",
        "date": "2026-05-06 + 2026-05-12",
        "signal": "Code with Claude conference — Managed Agents (Dreaming, Multiagent Orchestration, Outcomes, Webhooks) shipped as hosted service",
        "meaning": "Anthropic now ships an opinionated hosted agent runtime including memory consolidation, coordinator-subagent orchestration, rubric-graded outcomes, and HTTPS-signed webhooks — competing directly with in-house and open agent frameworks. Architects building bespoke agent stacks should justify build vs buy explicitly in 2H plans.",
        "source": "claude.com/blog/code-w-claude-sf-2026-sf"
      },
      {
        "vendor": "Alibaba",
        "date": "2026-05-11",
        "signal": "Qwen agentic shopping live across all 4B+ Taobao products",
        "meaning": "Largest production deployment of an LLM-driven conversational commerce surface to date — sets a real-world pattern for embedded agentic UX and gives Alibaba a data moat on conversational purchase signal ahead of the 618 festival. Western retail platforms should benchmark and respond within two quarters.",
        "source": "alibabacloud.com/blog/alibaba-opens-all-of-taobao-to-qwen-ai-ushering-in-a-new-agentic-shopping-experience"
      },
      {
        "vendor": "xAI",
        "date": "2026-05-14",
        "signal": "Grok Build coding-agent CLI launched, SuperGrok Heavy ($300/mo) tier",
        "meaning": "xAI enters the coding-agent CLI race (vs Claude Code, Codex, Cursor) with a local-first architecture and plan / review / approve model — the IDE / agent surface is now a competitive battleground; engineering managers should reassess developer tooling RFPs rather than assume current vendor lock-in holds through year-end.",
        "source": "ciodive.com/news/xAI-coding-agents-Grok-Build"
      },
      {
        "vendor": "DeepSeek",
        "date": "2026-05-09",
        "signal": "DeepSeek V4 image-recognition feature live; V4.1 scheduled June with full-modal coverage + MCP",
        "meaning": "DeepSeek closes the multimodal capability gap incrementally rather than waiting for a V4.1 step-change — and the June V4.1 commitment includes MCP support and enterprise toolchain, signaling a clear pivot from research demo to enterprise procurement target. Vendor managers should add DeepSeek to RFP shortlists rather than treating it as an experimental shadow.",
        "source": "cntechpost.com/2026/05/09/deepseek-roll-out-image-recognition-feature-ahead-v4-1-update"
      }
    ],
    "watchlist": [
      {
        "window": "May 19-20",
        "title": "Google I/O 2026 — Gemini 3.2 Flash expected launch (already leaked)",
        "why": "Gemini 3.2 Flash surfaced in iOS app and Eleuther Arena May 5 with ~92% of GPT-5.5 capability at ~$0.25 / $2.00 per Mtok; if confirmed at I/O, becomes the new price anchor for the cheap-frontier tier and forces OpenAI / Anthropic price moves within 60 days. Watch the fabric session for any second-hyperscaler MRC adoption."
      },
      {
        "window": "May 19-26",
        "title": "Anthropic / OpenAI counter-launch window post-I/O",
        "why": "Both Anthropic (Mythos full launch) and OpenAI (any 5.5.x refresh) historically respond to Google I/O within 7-10 days — high probability of a frontier-text move that will reset the LMArena top of the table; procurement contracts signed this week should leave headroom."
      },
      {
        "window": "May 19 - June 9",
        "title": "DeepSeek V4.1 — multimodal + MCP + enterprise toolchain",
        "why": "DeepSeek committed publicly to a June V4.1 release with full-modal (image, audio) coverage and Model Context Protocol support — this is the model that decides whether DeepSeek graduates to enterprise RFP shortlists or stays a research curiosity."
      },
      {
        "window": "May 19 - June 16",
        "title": "Kimi K3 (Moonshot) — 1M context, Kimi Linear attention",
        "why": "Prediction markets show ~74% probability of release inside ~6 weeks; rumored 3-4T parameters and Kimi Linear attention would put the next open frontier reasoning ceiling notably above current open SOTA — open-weights buyers should hold spend decisions if possible."
      },
      {
        "window": "May 19 - June 30",
        "title": "Meta MSL — Llama 5 or Muse Spark v2",
        "why": "Muse Spark (April 8) was MSL's first model and replaced the Llama frontier line; Meta has yet to ship a second MSL model or clarify Llama 5 status. Any move from Meta resets the open vs closed Meta strategy debate and the Llama ecosystem's roadmap; Llama-derivative shops should plan two scenarios."
      },
      {
        "window": "May 19 - June 30",
        "title": "Frontier voice response from Anthropic / Google",
        "why": "OpenAI's realtime trio with GPT-5-class reasoning (broadly available May 11) puts pressure on Anthropic and Google to ship a comparable voice-with-reasoning API or cede the call-center / accessibility / multilingual real-time segment for two quarters."
      }
    ],
    "changelog": [
      "First Sunday-cadence Pulse. Window shifts from Friday-publish to Sunday-publish (full ISO Mon-Sun) catching Saturday + early Sunday news.",
      "3 model rows added in W20 across 3 vendors (OpenBMB MiniCPM-V 4.6, Perceptron Mk1, NVIDIA SANA-WM). OpenAI realtime trio (W19) flipped to status: GA in updated.",
      "Scorecard reframed: specialist multimodal split (video / embodied + on-device) is now distinct from general multimodal — Edge / small added MiniCPM-V 4.6 challenger."
    ]
  }
}
