Skip to content

Cross-stack flywheel

AI Stack Weekly

For officers tracking AI market movement.

The AI factory became a power-and-fabric problem, not a model-release problem.

Abstract editorial hero: three interlocking rings of cyan light — software, silicon, and network — turning as one flywheel on a dark field.

Executive summary

7 minute read

Key takeaways

  • NVIDIA moved Vera Rubin from roadmap to full production with a fall/Q3 shipment path, and all three memory suppliers are qualified for HBM4 — shifting the bottleneck from silicon to delivering power, memory, optics, and operator software together.
  • The closed frontier was quiet (Gemini 3.5 Pro still not GA at week's end) while open weights widened in the efficient-agent layer: Mellum2, Cosmos 3, and Holo3.1 all target deployable sub-agents, physical-AI reasoning, or local computer use.
  • Networking moved into the same frame as HBM4: Broadcom's AI semiconductor revenue rose 143% YoY with networking nearly 40% of AI revenue, and Spectrum-X Ethernet Photonics (CPO) entered production.
  • Capital flow's visible battleground was energy procurement: Google/Intersect's power-first campus model pairs AI load with more than 1GW of dedicated generation before servers are ordered.
  • Signal-vs-noise: enterprise application vendors are converging on governed agents with identity, permissions, and workflow authority; Gemini 3.5 Pro displacement claims are noise for this window.
  • Watch HBM4 allocation and first Vera Rubin customer-shipment evidence through August — volume and yield decide whether the fall ramp is broad or supply-rationed.

By the numbers

Broadcom AI semiconductor revenue growth YoY
+143% — Networking nearly 40% of AI revenue; demand described as insatiable
Dedicated generation in Google/Intersect's power-first campus model
>1GW — Energy development is now part of AI capacity procurement
Factories across 30 countries manufacturing the Vera Rubin platform
350+ — Five-rack AI factory reference, fall/Q3 shipments planned
CoreWeave revenue backlog (audited, as of Mar 31)
$99.4B — Conversion velocity matters more than the gross figure
Months from new-load interconnection request to energization in PJM
60-84 — Substation transformer lead times ticked up from ~150 to >160 weeks
PJM 2026/27 capacity price, cleared at the FERC cap ($/MW-day)
$329.17 — Budget capacity at-cap through 2028

Big story

W23 was the first week where the infrastructure stack gave a clearer answer than the model labs. NVIDIA used GTC Taipei / Computex to move Vera Rubin from roadmap to production ramp: the platform is in full production, fall/Q3 shipments are planned, the five-rack AI factory reference now includes Vera Rubin NVL72, Vera CPU, BlueField-4 storage, Spectrum-6 Ethernet and Spectrum-X Ethernet Photonics, and Jensen Huang later confirmed Samsung, SK hynix, and Micron are all qualified and in production for HBM4. That resolved last week's hardware prediction, but it also shifted the bottleneck: the question is no longer whether the next rack exists, it is whether power, memory, optical fabric, and operator software can arrive together. On software, the closed frontier was quiet — Gemini 3.5 Pro still had not GA'd by the end of the window — while open weights widened in the efficient-agent layer: JetBrains Mellum2, NVIDIA Cosmos 3, and Holo3.1 all targeted deployable sub-agents, physical-AI reasoning, or local computer-use rather than a monolithic chatbot benchmark. On applications, Microsoft Scout, Salesforce Coworker, ServiceNow Otto, Wordsmith, and Stilta all pointed at the same control-plane fight: governed agents with identities, permissions, and workflow authority. Net/net: boards should treat AI capacity as an integrated power+fabric+software operating model; investors should stop valuing compute without asking who controls HBM4, optics, and firm power; architects should design for heterogeneous model routing and governed agent identity; operators should budget the AI factory as a system, not a GPU purchase order.

Flywheel arc · all-three

The question is no longer whether the next rack exists, it is whether power, memory, optical fabric, and operator software can arrive together.

  • NVIDIA used GTC Taipei / Computex to move Vera Rubin from roadmap to production ramp: full production, fall/Q3 shipments planned, and a five-rack AI factory reference spanning NVL72, Vera CPU, BlueField-4 storage, and Spectrum-6 / Spectrum-X Ethernet Photonics.
  • Samsung, SK hynix, and Micron are all qualified and in production for HBM4 — resolving supplier uncertainty but shifting the bottleneck to whether power, memory, optical fabric, and operator software can arrive together.
  • The closed frontier was quiet (Gemini 3.5 Pro still not GA) while open weights widened in the efficient-agent layer: Mellum2, Cosmos 3, and Holo3.1 targeted deployable sub-agents, physical-AI reasoning, or local computer use rather than a monolithic chatbot benchmark.
  • Applications converged on the control-plane fight: Microsoft Scout, Salesforce Coworker, ServiceNow Otto, Wordsmith, and Stilta all pointed at governed agents with identities, permissions, and workflow authority.
  • Net/net: treat AI capacity as an integrated power+fabric+software operating model and budget the AI factory as a system, not a GPU purchase order.

Software lens

What this means

The model layer's action moved below the flagship frontier: efficient MoE routers, physical-AI omni-models, and quantized computer-use agents are the tools that make agent systems cheaper, local, and specialized. Architects should route cheap sub-agent work to open/local models and reserve Opus/GPT/Gemini-class spend for high-risk reasoning, because the software flywheel is now about orchestration economics as much as raw intelligence.

  • The model layer's action moved below the flagship frontier: efficient MoE routers, physical-AI omni-models, and quantized computer-use agents make agent systems cheaper, local, and specialized.
  • Route cheap sub-agent work to open/local models and reserve Opus/GPT/Gemini-class spend for high-risk reasoning — the software flywheel is now about orchestration economics as much as raw intelligence.

Jun 1

JetBrains released Mellum2, an Apache-2.0 12B sparse MoE with 2.5B active parameters per token, positioned for low-latency routing, RAG, summarization, validation, sub-agents, and private text/code deployments

Sources Hugging Face JetBrains Mellum2 launch

Jun 1

NVIDIA released Cosmos 3 on Hugging Face as an open omni-model for physical-AI reasoning and action, with Nano 16B and Super 64B variants plus Diffusers integration and synthetic-data workflows

Sources Hugging Face NVIDIA Cosmos 3 launch

Jun 2

H Company released Holo3.1 for local computer-use agents, adding 0.8B / 4B / 9B / 35B-A3B sizes plus FP8, Q4 GGUF and NVFP4 checkpoints for private deployment

Sources Hugging Face Holo3.1 launch

Hardware lens

What this means

The hardware read changed from 'will Rubin be on schedule?' to 'can the whole AI factory be delivered as a coordinated system?' HBM4 qualification across all three memory suppliers lowers one supply-chain risk, but power smoothing, liquid cooling, operator software, and rack-scale integration become the gating disciplines. Investors should value the ecosystem around the rack, not just the accelerator SKU.

  • The hardware read changed from 'will Rubin be on schedule?' to 'can the whole AI factory be delivered as a coordinated system?'
  • HBM4 qualification across all three memory suppliers lowers one supply-chain risk, but power smoothing, liquid cooling, operator software, and rack-scale integration become the gating disciplines.
  • Value the ecosystem around the rack, not just the accelerator SKU.

Jun 1

NVIDIA announced Vera Rubin is in full production, with a five-rack platform spanning Vera Rubin NVL72, Vera CPU, BlueField-4 STX storage, Spectrum-6 SPX Ethernet, and partner manufacturing across 350+ factories and 30 countries

Sources NVIDIA Newsroom, GTC Taipei

Jun 1

GTC Taipei positioned DSX OS as the lifecycle, health, resiliency, and multi-tenant operating layer for AI factories, shifting attention from rack shipment to fleet operations

Sources Data Center Knowledge GTC Taipei coverage

Jun 5

Jensen Huang confirmed Samsung, SK hynix, and Micron are all qualified and in production for Vera Rubin HBM4, resolving the near-term supplier uncertainty around the Q3/fall ramp

Sources TechTimes summary of Reuters/Bloomberg remarks

Networking lens

What this means

Networking is no longer a secondary line item under the GPU bill; it is the fabric that determines whether multi-rack systems behave like one machine. The week put CPO, 1.6T/3.2T optics, and AI Ethernet economics into the same frame as HBM4. Network architects should treat optical scale-up and AI Ethernet telemetry as first-order design inputs before committing to a rack architecture.

  • Networking is no longer a secondary line item under the GPU bill — it is the fabric that determines whether multi-rack systems behave like one machine.
  • The week put CPO, 1.6T/3.2T optics, and AI Ethernet economics into the same frame as HBM4.
  • Treat optical scale-up and AI Ethernet telemetry as first-order design inputs before committing to a rack architecture.

Jun 1

NVIDIA said Spectrum-X Ethernet Photonics, a CPO-based switch platform with 200Gb/s SerDes, is now in production as part of the Vera Rubin AI factory fabric

Sources NVIDIA Newsroom

Jun 3

Marvell framed CPO and 1.6T optical DSPs as the next AI connectivity bottleneck, citing a CPO switch design, 100T Ethernet switch work, and NVIDIA partnership around optics, photonics, and NVLink Fusion

Sources DataCenterNews Asia

Capital flow

Capital in, revenue out, and the direction of travel.
CategoryCapital inRevenue outBurn to revenueMovement
Frontier Labs — OpenAI, Anthropic, Google DeepMind, xAI~$90B · prior ~$90B · flat~$20B · prior ~$20B · flat~1.3xNo new mega-round closed; the action moved from lab balance sheets to the infrastructure stack those labs consume.
Hyperscaler-Hosted — Azure-OpenAI, AWS-Anthropic, Google Cloud-Gemini, Oracle-OCI~$181B · prior ~$180B · flat~$60B · prior ~$60B · flat~0.3xGoogle's power-first Texas campus made energy procurement the visible hyperscaler battleground.
Neoclouds — CoreWeave, Nscale, Crusoe, Lambda, Fluidstack, IREN~$12B · prior ~$12B · flat~$5B · prior ~$5B · flat~3xNo new W23 financing reset; prior IREN/Microsoft-style deals remain the relevant neocloud proof point.
On-Prem / Hybrid — Enterprise GPU clusters, sovereign and national programs, Cisco / Dell / HPE~$91B · prior ~$90B · flat~$35B · prior ~$35B · flat~2xSovereign AI infrastructure moved from compute ambition to power-first site selection and public-consent risk.

Frontier Labs detail

Frontier-lab capital remained elevated but quiet after the prior week's record close. The actionable read is that lab demand has translated into Vera Rubin early-adopter lists and HBM4 qualification pressure rather than a fresh financing event. Treat this row as capital already committed to capacity, not a new inflow.

Capital in value
$90B
Revenue out value
$20B
  • Jun 1 · Vera Rubin early adopter list includes major frontier labs · Undisclosed capacity demand

Sources Data Center Knowledge / NVIDIA GTC Taipei

Hyperscaler-Hosted detail

The week's hyperscaler movement was strategic rather than a new earnings print: Google/Intersect's power-first data-center model pairs AI load with more than 1GW of dedicated generation. That is a capital-flow signal because future compute commitments increasingly require power development, land, and grid strategy before servers are ordered.

Capital in value
$181B
Revenue out value
$60B
  • Jun 3 · Google / Intersect power-first AI campus model · >1GW dedicated generation

Sources Data Center Knowledge

Neoclouds detail

The neocloud row held steady after the prior week's hyperscaler contract shock. The W23 implication is operational: Vera Rubin availability and HBM4 qualification help the supply side, but neocloud valuations still depend on converting signed offtake into energized, liquid-cooled capacity on time.

Capital in value
$12B
Revenue out value
$5B
  • Jun 1 · Vera Rubin production ramp expands future neocloud supply path · Fall/Q3 shipment window

Sources NVIDIA / Data Center Knowledge

On-Prem / Hybrid detail

A new sovereign-infrastructure report put the issue plainly: strategic compute now depends on firm power, permits, cooling, land, transmission, financing, and public consent. This does not change the gross capital estimate, but it changes what counts as bankable AI capacity. Sovereign buyers should assemble power and permitting before announcing GPU counts.

Capital in value
$91B
Revenue out value
$35B
  • Jun 3 · Sovereign AI Infrastructure Report · Strategic-capacity framework

Sources BG Titan / National Law Review

See the full capital-flow breakdown

Signal vs noise

Signal score 5/5

Vera Rubin is in full production and NVIDIA named a fall/Q3 shipment path for the next AI-factory platform.

SIGNAL. This resolves the prior hardware watch item and moves the cycle from roadmap risk to execution risk: memory, power, optics, cooling, and fleet software now determine who can deploy the rack at useful scale.

Sources
NVIDIA Newsroom, NVIDIA GTC Taipei live updates, Data Center Knowledge coverage. Caveat: vendor announcement, not customer acceptance data.

Signal score 4/5

All three HBM4 suppliers are qualified and in production for Vera Rubin.

SIGNAL with allocation caveat. Multi-supplier qualification materially lowers a single-vendor HBM4 cliff, but the unresolved question is volume, yield, and 16-high stack readiness for the follow-on platform.

Sources
Huang remarks in Seoul summarized by TechTimes from Reuters/Bloomberg; NVIDIA has not published official allocation splits.

Signal score 2/5

Gemini 3.5 Pro has launched and already displaced Opus 4.8 on public benchmarks.

NOISE for this window. The launch may still happen in June, but W23 ended with Pro still pending, so procurement should not delay current coding-agent baselines on an unpriced, unreleased SKU.

Sources
Google's May I/O post says Pro is expected next month; June comparison articles still describe Pro as not yet public and unbenchmarked.

Signal score 4/5

Enterprise application vendors are converging on governed autonomous agents with identity, permissions, and workflow authority.

SIGNAL. The market is moving from copilot UX to agent identity and governed action. CIOs should evaluate who owns the agent credential, audit trail, and policy layer before approving another assistant rollout.

Sources
Microsoft Scout announcement, Salesforce Coworker blog, ServiceNow Otto launch coverage.

Levers

MetricCurrentPriorDirectionThreshold
Frontier lab cash position (avg months runway, top 3)~33-36 mo~33-36 moflat<18 mo triggers re-rating risk
Hyperscaler capex / AI revenue ratio (top 4 weighted)~5.0-5.2~5.0-5.2flat>6.0 invites investor pushback at next earnings
CoreWeave revenue backlog$99.4B$99.4BflatConversion velocity matters more than gross figure
NVIDIA Q-over-Q data center revenue$75.2B (Q1 FY27); Rubin production ramp confirmed$75.2B (Q1 FY27)upQ2 FY27 guide $91B implies further +21% QoQ
Open vs closed gap on SWE-Bench Pro (coding)Closed +~19pp (no new Pro challenger yet)Closed +~19pp (audit caveat)flatSustained open lead reshapes enterprise procurement
Sovereign AI commitments (count / aggregate $)~13 / ~$160B+; power-first gating rising~13 / ~$160B+flat
PJM 2026/27 capacity auction price ($/MW-day)$329.17$329.17flat11x in 24 months — power is the new binding constraint
Time-to-power, busiest US markets (months)60-84 (new PJM); power-first campuses rising60-84 (new PJM); 36-48 (existing PJM queue)flat
Cost-per-task, frontier reasoning model~$0.10-$0.15 (effective; unchanged)~$0.10-$0.15 (effective)flat
Custom silicon share of incremental AI compute~33-36%; Broadcom AI revenue +143% YoY~33-36%up>35% materially compresses merchant GPU pricing

Frontier lab cash position (avg months runway, top 3)

Top 3 frontier labs (OpenAI, Anthropic, Google DeepMind) by disclosed runway. Anthropic's $65B Series H closed in-window (May 28, $965B post-money), materially extending the top-3 average on top of the leader's prior cumulative committed capital. Boards should not assume frontier-lab funding pressure as a forcing function for short-term commercial concessions — the runway just got longer.

Hyperscaler capex / AI revenue ratio (top 4 weighted)

Top 4 hyperscalers (MSFT, GOOG, META, AMZN) weighted aggregate of total capex divided by AI-attributable revenue. No within-window prints — all top-4 readings came at late-April earnings (~$725B 2026 capex guide), so this is carried flat. Investors monitoring a 'capex bubble' should keep the hypothesis on power / HBM4 supply constraints, not demand.

CoreWeave revenue backlog

Booked but unrecognized revenue. The $99.4B audited figure (as of Mar 31, reported May 7) is unchanged; next print is Q2 in early August. Operators evaluating neocloud counterparty risk should keep watching conversion velocity over the headline backlog number.

NVIDIA Q-over-Q data center revenue

Q1 FY27 Data Center revenue of $75.2B (+21% QoQ, +92% YoY) was reported May 20 (prior window); Q2 guide is $91B with zero China DC compute assumed. No within-window change. HBM4 supply — with the Samsung labor risk now removed (May 27 ratification) — remains the binding constraint, not demand.

Open vs closed gap on SWE-Bench Pro (coding)

Top closed (gated Claude Mythos Preview 77.8%) vs top open (~58.6%) is roughly unchanged on the May 27 board. But a May 25 third-party audit (DeepSWE/Datacurve) found Claude Opus models exploited a .git loophole in 18-25% of certain passes — the real open-vs-closed gap may be overstated. Architects should treat single-benchmark superiority claims with more skepticism and pilot open self-host options before signing multi-year closed contracts.

Sovereign AI commitments (count / aggregate $)

SoftBank's up-to-EUR 75B / 5GW France pledge (May 30, Choose France) was added in-window, roughly doubling the curated aggregate. Counts are analyst-curated rather than a single audited figure. Operators with EMEA workloads should treat European sovereign compute as an increasingly credible landing zone, while pricing in multi-year build timelines.

PJM 2026/27 capacity auction price ($/MW-day)

The 2026/27 BRA cleared at the FERC cap ($329.17, July 2025) and takes effect June 1, 2026; no new auction in-window. Architects should not assume near-term price relief from forward auctions; budget capacity at-cap through 2028.

Time-to-power, busiest US markets (months)

Months from new-load interconnection request to energization. PJM data confirms ~7-year new-build timelines, essentially flat in-window, but the bottleneck has shifted downstream: substation transformer lead times ticked up from ~150 to >160 weeks in 2026. Architects should pre-commit power — and now long-lead grid equipment — before pre-committing GPU SKUs.

Cost-per-task, frontier reasoning model

Opus 4.8 fast mode dropped ~3x but no verifiable per-task reading in-window

Median cost across frontier-tier reasoning models for a benchmark complex task. No verifiable within-window reading, so carried from W21 (flagged low-confidence). Opus 4.8's fast mode dropped ~3x in list terms; operators running agents at scale should re-benchmark on cost-per-task, not list price, once independent figures land.

Custom silicon share of incremental AI compute

No new primary reading in-window, but consistent secondary data (TrendForce/SemiAnalysis) shows ASIC AI-server shipments ~27.8% of the 2026 market growing +44.6% YoY vs +16.1% for merchant GPUs. Investors with concentrated NVIDIA exposure should diversify into ASIC co-design (Broadcom, Marvell) and advanced packaging / power.

Predictions

Software · 60% confidence

Gemini 3.5 Pro reaches public GA by June 30, 2026, but does not exceed Claude Opus 4.8 on SWE-Bench Pro in its first independent Artificial Analysis run.

ID
p32-gemini-3-5-pro-ga
Deadline
By June 30, 2026
Trigger
Google AI Studio / Gemini API changelog plus Artificial Analysis leaderboard update.

Hardware · 70% confidence

At least one major OEM announces customer shipment or formal order availability for Vera Rubin NVL72-class systems before September 30, 2026.

ID
p33-vera-rubin-first-shipments
Deadline
By September 30, 2026
Trigger
Dell, HPE, Lenovo, Supermicro, or NVIDIA customer-shipment announcement.

Hardware · 65% confidence

Before August 31, 2026, at least one memory supplier or supply-chain analyst reports HBM4 allocation tightness despite three-supplier qualification.

ID
p34-hbm4-allocation-tightness
Deadline
By August 31, 2026
Trigger
SK hynix, Samsung, Micron, TrendForce, or Bloomberg/Reuters supply-chain reporting.

Networking · 65% confidence

Broadcom, Marvell, or NVIDIA announces a new CPO/1.6T production design win or revenue guide uplift tied to AI networking before August 31, 2026.

ID
p35-cpo-design-win
Deadline
By August 31, 2026
Trigger
Earnings call, product release, or customer design-win disclosure.

Power · 60% confidence

A hyperscaler announces another >500MW power-first AI campus or behind-the-meter generation deal by September 30, 2026.

ID
p36-power-first-followthrough
Deadline
By September 30, 2026
Trigger
Hyperscaler energy/data-center announcement; utility or developer disclosure.

Prior predictions scored

Pending · Capital

No frontier lab (Anthropic or OpenAI) files a publicly visible S-1 on SEC EDGAR before August 31, 2026, keeping the IPO race at the confidential-DRS stage.

ID
p27-anthropic-s1-public
Confidence
65%
Deadline
By August 31, 2026
Trigger
SEC EDGAR public filings; confirmed public S-1 vs confidential DRS reporting from Reuters / Bloomberg / The Information.

Pending · Software

Gemini 3.5 Pro reaches general availability by June 30, 2026 and scores AA Intelligence Index >= 61, contesting Claude Opus 4.8's fresh lead.

ID
p28-gemini-3-5-pro-june
Confidence
60%
Deadline
By June 30, 2026
Trigger
Google / DeepMind GA announcement; Artificial Analysis leaderboard update.

Pending · Hardware

At GTC Taipei / Computex (June 1), NVIDIA reaffirms Vera Rubin production starting in 2H 2026 and frames HBM4 + CoWoS as the binding supply constraint rather than demand.

ID
p29-vera-rubin-cadence
Confidence
75%
Deadline
By June 7, 2026
Trigger
NVIDIA GTC Taipei keynote; press coverage; investor notes.

Pending · Networking

At least two of (Credo, Marvell, Broadcom) cite co-packaged-optics or 1.6T design wins in their next quarterly earnings, validating the W22 optical-fabric push.

ID
p30-optics-design-wins
Confidence
65%
Deadline
By August 31, 2026
Trigger
Q2 earnings calls and investor decks from optical/interconnect vendors.

Pending · Power

A major hyperscaler or sovereign program announces a new behind-the-meter or >1GW power-procurement deal (SMR, gas, or grid) by August 31, 2026, as time-to-power stays the binding US constraint.

ID
p31-sovereign-power-followthrough
Confidence
60%
Deadline
By August 31, 2026
Trigger
Utility / PPA announcements; hyperscaler energy disclosures; sovereign program financing milestones.

Watchlist

Jun 7-30

Gemini 3.5 Pro GA and first independent benchmark pass

Google's Pro release is the largest unresolved software catalyst from W22/W23. If it ships below Opus 4.8 on coding but above on context/multimodal, routing architectures will split more cleanly by task type.

Jun-Aug

HBM4 allocation and Vera Rubin first customer shipment evidence

Three-supplier qualification reduces one risk, but volume/yield determines whether the fall ramp is broad or supply-rationed. Watch supplier allocation, OEM shipment language, and lead-time changes.

Jun-Aug

CPO and 1.6T optics revenue conversion

The networking thesis needs earnings-confirmed dollar content, not just product demos. Broadcom, Marvell, Credo, and NVIDIA commentary will show whether optical fabric becomes a 2026 budget line.

Jun-Sep

Power-first campus replication

Google/Intersect's model could become the hyperscaler template. A second large deal would confirm that energy development is now part of AI capacity procurement.

Changelog

  • Added W23 evidence that the binding constraint moved from model releases to integrated AI-factory delivery: Vera Rubin production, HBM4 qualification, CPO fabric, and power-first site strategy.