---
title: >-
  The harness became the product: OpenAI and Anthropic shipped rival work runtimes 48 hours apart —
  while memory and power marked the scarcity trade to market with $26.5B of real money.
publication: The AI Stack Weekly
slug: 2026-W28
issueNumber: 12
isoYear: 2026
isoWeek: 28
publishedAt: '2026-07-11'
canonicalUrl: https://brianletort.ai/industry/weekly/2026-W28
pdfUrl: https://brianletort.ai/downloads/ai-stack-weekly-2026-W28.pdf
schemaVersion: 2026.05.02
flywheelArc: all-three
capitalFlow:
  - category: Frontier Labs
    capitalIn: ~$95B
    capitalInPrior: ~$95B
    capitalInDirection: flat
    revenueOut: ~$21B
    revenueOutPrior: ~$21B
    revenueOutDirection: flat
    burnToRevenue: ~1.3x
  - category: Hyperscaler-Hosted
    capitalIn: ~$187B
    capitalInPrior: ~$187B
    capitalInDirection: flat
    revenueOut: ~$62B
    revenueOutPrior: ~$62B
    revenueOutDirection: flat
    burnToRevenue: ~3.0x
  - category: Neoclouds
    capitalIn: ~$13.5B
    capitalInPrior: ~$13.5B
    capitalInDirection: flat
    revenueOut: ~$5B
    revenueOutPrior: ~$5B
    revenueOutDirection: flat
    burnToRevenue: ~2.7x
  - category: On-Prem / Hybrid
    capitalIn: ~$94.5B
    capitalInPrior: ~$94B
    capitalInDirection: up
    revenueOut: ~$36B
    revenueOutPrior: ~$36B
    revenueOutDirection: flat
    burnToRevenue: ~2.6x
levers:
  - metric: Frontier lab cash position (avg months runway, top 3)
    current: >-
      ~34-37 mo; Anthropic at implied $1.2T on secondaries (broker-reported), IPO calendars
      unchanged
    prior: ~34-37 mo; OpenAI leaning 2027 IPO, Anthropic holding Oct 2026
    direction: flat
    threshold: <18 mo triggers re-rating risk
  - metric: Hyperscaler capex / AI revenue ratio (top 4 weighted)
    current: ~5.0-5.3; Meta targets 14 GW of compute in 2027 (2x 2026) on $125-145B capex guidance
    prior: ~5.0-5.3; free-cash-flow crossover framed for ~Q3 2026
    direction: flat
    threshold: '>6.0 invites investor pushback at next earnings'
  - metric: CoreWeave revenue backlog
    current: >-
      ~$100B reported; Helios Phase I (133 MW) delivered on schedule — backlog now converting to
      lease revenue
    prior: ~$100B reported; Meta Compute repriced concentration risk (stock -12-15% Jul 1)
    direction: flat
    threshold: Conversion velocity matters more than gross figure
  - metric: NVIDIA Q-over-Q data center revenue
    current: >-
      $75.2B Q1 FY27; Q2 guide $91B (reports Aug 26); SemiAnalysis sees H2 ~20% above consensus
      despite Kyber dispute
    prior: $75.2B Q1 FY27; Q2 guide $91B, reports Aug 26
    direction: flat
    threshold: Q2 FY27 guide $91B implies further +21% QoQ
  - metric: Open vs closed gap on coding (SWE-Bench / agentic)
    current: >-
      Effectively closed on cost-quality: GLM 5.2 statistically tied with Opus 4.8 at $1.28 vs
      $1.94/task (Databricks); Hy3 adds Apache-2.0 agentic-search lead
    prior: >-
      Narrowing from both sides: closed prices down (Sonnet 5 $2/$10; Terra promised at half
      GPT-5.5), open pressure sustained
    direction: down
    threshold: Sustained open lead reshapes enterprise procurement
  - metric: Sovereign AI commitments (count / aggregate $)
    current: >-
      ~14 / ~$180B+ (flat; Meta Alberta and MARA Texas are corporate capital on power-rich land, not
      sovereign programs)
    prior: ~14 / ~$180B+ (flat; SB Neo's 10GW is corporate, not sovereign)
    direction: flat
    threshold: null
  - metric: PJM 2026/27 capacity auction price ($/MW-day)
    current: >-
      $329.17; 2028/29 BRA bids closed Jul 7 — results post Jul 14 after 4 p.m. ET (prediction p54
      resolves)
    prior: $329.17; 2028/29 BRA bids close Jul 7, results Jul 14 (slipped ~1 week)
    direction: flat
    threshold: 11x in 24 months — power is the new binding constraint
  - metric: Time-to-power, busiest US markets (months)
    current: >-
      60-84; hyperscalers routing around queues — Meta fully funds its own generation in Alberta
      because the grid cannot host multiple large loads
    prior: 60-84; FERC intervenor deadline Jul 9, tariff responses due Aug 17
    direction: flat
    threshold: null
  - metric: Cost-per-task, frontier reasoning model
    current: >-
      ~$0.06-$0.12 effective; GPT-5.6 GA tiering (Sol $5/$30 / Terra $2.50/$15 / Luna $1/$6), Grok
      4.5 at $2/$6 — and harness choice swings per-task cost 2x+
    prior: >-
      ~$0.08-$0.13 effective; Sonnet 5 $2/$10 intro, GPT-5.6 tiering Sol $5/$30 / Terra $2.50/$15 /
      Luna $1/$6
    direction: down
    threshold: null
  - metric: Custom silicon share of incremental AI compute
    current: >-
      ~34-37%; Meta's Iris enters production in September, Broadcom-Apple extended through 2031
      (8-K), AWS raising Trainium 3 orders 20-30%
    prior: ~33-36%; silicon-IP layer forming (Oxmiq) to lower custom-ASIC entry cost
    direction: up
    threshold: '>35% materially compresses merchant GPU pricing'
predictions:
  - id: p57-gemini-3-5-pro-ga-jul31
    lens: software
    confidencePct: 58
    deadline: By July 31, 2026
    text: >-
      Gemini 3.5 Pro reaches public general availability — a callable API model ID with published
      pricing — by July 31, 2026, after slipping past its June window and the reported July 17
      target.
  - id: p58-harness-cost-telemetry
    lens: software
    confidencePct: 64
    deadline: By August 31, 2026
    text: >-
      At least one major agent platform (OpenAI, Anthropic, GitHub, or Cursor) ships product-level
      per-task or per-harness cost telemetry or routing controls — beyond session budget caps — by
      August 31, 2026.
  - id: p59-tsmc-q2-capex-raise
    lens: hardware
    confidencePct: 62
    deadline: By July 16, 2026
    text: >-
      TSMC's July 16 Q2 earnings raise or reiterate the top end of full-year 2026 capex guidance and
      report HPC/AI platform revenue up more than 50% year over year, confirming the
      packaging-constrained AI capex ramp.
  - id: p60-scale-across-follow-on
    lens: networking
    confidencePct: 61
    deadline: By September 30, 2026
    text: >-
      A second named vendor or operator announces a commercial cross-data-center scale-across AI
      fabric deployment or product launch — following DriveNets/WhiteFiber — by September 30, 2026.
predictionsPrior:
  - id: p53-skhy-debut-validates-memory
    lens: capital
    outcome: partial
    deadline: By July 31, 2026
    text: >-
      SK hynix's Nasdaq ADS offering prices at or above its indicated ~$166/ADS level and closes its
      first trading week above the offer price, by July 31, 2026.
  - id: p54-pjm-2028-29-at-cap
    lens: power
    outcome: pending
    deadline: By July 14, 2026
    text: >-
      The PJM 2028/29 base residual auction clears within 5% of the ~$325/MW-day cap when results
      post on July 14, 2026.
  - id: p55-terra-confirms-repricing-cycle
    lens: software
    outcome: hit
    deadline: By August 31, 2026
    text: >-
      GPT-5.6 reaches broad GA with the Terra tier priced at or below $2.50/$15 per MTok — half of
      GPT-5.5's rate — confirming a closed-lab repricing cycle rather than a one-off Sonnet 5 cut,
      by August 31, 2026.
  - id: p56-samsung-hbm4-to-nvidia
    lens: hardware
    outcome: hit
    deadline: By August 31, 2026
    text: >-
      Samsung's HBM4 supply to NVIDIA is publicly confirmed — via earnings call, company statement,
      or multi-source supply-chain reporting — by August 31, 2026.
signalScores:
  - 5
  - 4
  - 3
  - 2
  - 1
keyTakeaways:
  - >-
    The harness became the product: Anthropic pushed Claude Cowork to web and mobile Jul 7, and
    OpenAI answered Jul 9 with ChatGPT Work — the Codex task runtime generalized to all knowledge
    work, bundled into every plan including Free.
  - >-
    Two independent studies showed harness choice now swings agent cost more than model choice —
    over 2x per task at equal quality (Databricks) and ~10x via harness tuning (LangChain/NVIDIA);
    benchmark the harness, not the rate card.
  - >-
    The model layer kept commoditizing on cue: GPT-5.6 went GA with Terra at $2.50/$15 — half of
    GPT-5.5's rate, resolving prediction p55 as a hit — and Grok 4.5 launched at $2/$6 positioning
    explicitly on cost-per-task.
  - >-
    Memory and power marked the scarcity trade to market with real money: SK hynix closed the
    largest-ever foreign US IPO at $26.5B with 7x demand, and Samsung guided to a record ~KRW 89.4T
    quarter on AI memory.
  - >-
    Meta committed C$13B to a 1 GW Alberta campus where it must fund its own gas generation because
    the grid cannot host multiple large AI loads — the hyperscaler-as-utility template.
  - >-
    Watch Jul 14: PJM's 2028/29 capacity auction results resolve prediction p54 — the week's
    highest-information event on whether power stays the binding constraint through 2028.
byTheNumbers:
  - value: $26.5B
    label: SK hynix's close on the largest-ever foreign US IPO
  - value: $2.50/$15
    label: GPT-5.6 Terra per MTok — exactly half of GPT-5.5's rate
  - value: ~KRW 89.4T
    label: Samsung's record Q2 operating profit guidance, +1,810% YoY
  - value: 2x+
    label: Per-task cost swing from harness choice alone, at equal quality
  - value: C$13B
    label: Meta's 1 GW Alberta campus — its largest outside the US
  - value: ~30 hours
    label: Tencent's Apache-2.0 Hy3 from release to local deployment
---

# The harness became the product: OpenAI and Anthropic shipped rival work runtimes 48 hours apart — while memory and power marked the scarcity trade to market with $26.5B of real money.

*Issue 12 · Week 28 of 2026 · Published 2026-07-11*

## Executive summary

- The harness became the product: Anthropic pushed Claude Cowork to web and mobile Jul 7, and OpenAI answered Jul 9 with ChatGPT Work — the Codex task runtime generalized to all knowledge work, bundled into every plan including Free.
- Two independent studies showed harness choice now swings agent cost more than model choice — over 2x per task at equal quality (Databricks) and ~10x via harness tuning (LangChain/NVIDIA); benchmark the harness, not the rate card.
- The model layer kept commoditizing on cue: GPT-5.6 went GA with Terra at $2.50/$15 — half of GPT-5.5's rate, resolving prediction p55 as a hit — and Grok 4.5 launched at $2/$6 positioning explicitly on cost-per-task.
- Memory and power marked the scarcity trade to market with real money: SK hynix closed the largest-ever foreign US IPO at $26.5B with 7x demand, and Samsung guided to a record ~KRW 89.4T quarter on AI memory.
- Meta committed C$13B to a 1 GW Alberta campus where it must fund its own gas generation because the grid cannot host multiple large AI loads — the hyperscaler-as-utility template.
- Watch Jul 14: PJM's 2028/29 capacity auction results resolve prediction p54 — the week's highest-information event on whether power stays the binding constraint through 2028.

**By the numbers.**

- **$26.5B** — SK hynix's close on the largest-ever foreign US IPO (Priced at $149/ADS but 7x oversubscribed, closing day one up ~13%)
- **$2.50/$15** — GPT-5.6 Terra per MTok — exactly half of GPT-5.5's rate (Prediction p55 resolves as a hit seven weeks early)
- **~KRW 89.4T** — Samsung's record Q2 operating profit guidance, +1,810% YoY (Reportedly the largest quarterly operating profit ever posted by a tech company)
- **2x+** — Per-task cost swing from harness choice alone, at equal quality (Databricks merged-PR benchmark; LangChain/NVIDIA showed ~10x via harness tuning)
- **C$13B** — Meta's 1 GW Alberta campus — its largest outside the US (Self-funded 932 MW gas tolling deal because the grid cannot host multiple large loads)
- **~30 hours** — Tencent's Apache-2.0 Hy3 from release to local deployment (llama.cpp MTP speculative decoding measured +40% throughput)

## Big Story

Two launches this week redrew the enterprise AI procurement map. Anthropic pushed Claude Cowork to web and mobile on July 7 with cloud-run background sessions, citing 1.2 million sessions across 600,000+ organizations showing most Cowork use is non-coding knowledge work. OpenAI answered on July 9 with ChatGPT Work — the Codex task runtime generalized to all knowledge work, bundled into a desktop app available on every plan including Free, with 1 million of Codex's 5 million weekly users already working outside software development. The same week, two independent quantitative studies explained why the harness is where the fight moved: Databricks' merged-PR benchmark found the same model at the same effort costs over 2x more per task depending on harness choice — with open-weight GLM 5.2 statistically tied with Opus 4.8 at $1.28 vs $1.94 per task — and LangChain/NVIDIA showed harness tuning alone lifts an open model to near-Opus quality at roughly 10x lower cost. The model layer kept commoditizing on cue: GPT-5.6 went GA with Terra at $2.50/$15 (half of GPT-5.5's rate, resolving prediction p55 as a hit), and Grok 4.5 launched at $2/$6 positioning explicitly on cost-per-task. Meanwhile the physical layer banked the other side of the trade: SK hynix closed the largest-ever foreign US IPO at $26.5B with 7x demand and a +13% first-day pop, Samsung guided to a record ~KRW 89.4T quarter on AI memory, and Meta committed C$13B to a 1 GW Alberta campus where it must fund its own gas generation because the grid cannot host multiple large AI loads. Net/net: capability is commoditizing, harnesses are consolidating into suites, and margin keeps pooling in memory and power. Boards should treat agent-suite governance (defaults, budgets, audit) as this quarter's control gap; investors should note the scarcity trade just got public-market confirmation; architects should re-run agent cost benchmarks at the harness level, not the rate card.

Flywheel arc: `all-three`.

## Software lens

- **Jul 9.** GPT-5.6 reached full GA as a three-tier family — Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per MTok — after a 12-day government-coordinated preview; GPT-5.4 retires Jul 23 _([OpenAI; Vellum benchmark analysis](https://openai.com/index/gpt-5-6/))_
- **Jul 8.** xAI released Grok 4.5, an 'Opus-class' workhorse at $2/$6 per MTok co-trained with Cursor — 4th on the AA Intelligence Index, claiming ~4.2x fewer output tokens per SWE-Bench Pro task; not available in the EU at launch _([xAI; TechCrunch](https://x.ai/news/grok-4-5))_
- **Jul 6-9.** Tencent released Hy3 (295B MoE, 21B active) under clean Apache 2.0 at ~$0.20/$0.80 per MTok — and the community made it locally deployable in ~30 hours, with llama.cpp MTP speculative decoding measuring +40% throughput _([Hugging Face model card; Simon Willison; llama.cpp PR #25395](https://huggingface.co/tencent/Hy3))_
- **Jul 9.** Meta released Muse Spark 1.1 and opened the Meta Model API in public preview — the first frontier Meta model distributed through a first-party API rather than open weights or Meta's own apps _([Meta AI](https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/))_
- **Jul 8.** Gemini 3.5 Pro slipped again to a reported July 17 target after Google rebuilt the base model; the public API still lists no 3.5 Pro model ID _([TechTimes; Google Gemini API model list](https://ai.google.dev/gemini-api/docs/models))_

**What this means.** Four frontier-relevant releases in four days, and every closed launch priced against the open floor: Terra at half GPT-5.5's rate and Grok 4.5 at $2/$6 are responses to Apache-2.0 Hy3 economics and GLM 5.2's per-task parity. Architects should lock inference pricing during this window and re-benchmark at the task level — see the Model Pulse for the full architecture read, including why the harness now matters more than the rate card.

## Hardware lens

- **Jul 7.** Samsung guided to a record ~KRW 89.4T Q2 operating profit (+1,810% YoY) on AI memory — reportedly the largest quarterly operating profit ever posted by a tech company, with HBM4 reaching $1B in sales within four months _([Samsung Newsroom; Seoul Economic Daily; Korea Herald](https://news.samsung.com/global/samsung-electronics-announces-earnings-guidance-for-second-quarter-2026))_
- **Jul 10.** SK hynix closed the largest-ever foreign US IPO at $26.5B — priced at $149/ADS (below the ~$166 indication) but 7x oversubscribed, closing day one up ~13%; proceeds fund the Yongin fab, HBM packaging, and EUV tools _([TechCrunch; Korea Herald; Yahoo Finance](https://techcrunch.com/2026/07/10/sk-hynix-raises-26-5b-in-the-biggest-foreign-ipo-in-us-history-is-urged-to-build-new-us-fabs/))_
- **Jul 6.** SemiAnalysis reported NVIDIA's ~600 kW Kyber rack for Rubin Ultra slipping to 2028 on PCB-midplane manufacturability; NVIDIA publicly denied it the same day, holding to Kyber racks in H2 2027 _([CNBC; Wccftech (NVIDIA statement)](https://www.cnbc.com/2026/07/06/nvidia-kyber-rack-system-delays-manufacturing-taiwan-rubin-chips-.html))_
- **Jul 9.** Micron raised planned US investment to $250B+ through 2035 and committed up to $3B to the domestic supply chain, including $500M in GlobalWafers' Texas 300mm plant — with its HBM sold out for 2026 and able to fill only 50-66% of demand _([Micron press releases; Tom's Hardware](https://www.globenewswire.com/news-release/2026/07/09/3324807/14450/en/Micron-Accelerates-U-S-Investments-Pours-First-Concrete-at-New-York-Fab.html))_
- **Jul 9.** Reuters: Meta's Broadcom-designed 'Iris' MTIA chip enters production in September, on a roadmap of one new chip every ~6 months as Meta targets doubling compute to 14 GW in 2027 _([Reuters; TechCrunch; DCD](https://techcrunch.com/2026/07/09/metas-new-ai-chips-will-begin-production-in-september/))_

**What this means.** Memory completed its move from allocation story to capital-markets story — a record quarter, a $26.5B IPO with 7x demand, and TrendForce showing long-term agreements now capping price increases means the scarcity is being contractually locked, not loosening. Operators should treat the SemiAnalysis-vs-NVIDIA Kyber dispute as a live facility-planning risk (600 kW racks slipping would reshape 2027-28 datacenter designs) and watch Samsung's Jul 30 divisional print to make the HBM4-to-NVIDIA confirmation unambiguous.

## Networking lens

- **Jul 9.** DriveNets and WhiteFiber deployed the first commercial long-distance scale-across AI supercluster — two H200 sites 83 km apart validated at 111.2 Tbps with 0.9 ms guaranteed latency, within 8% of the physical limit of light in fiber _([DriveNets; WhiteFiber; Light Reading](https://www.prnewswire.com/il/news-releases/drivenets-announces-industrys-first-commercial-deployment-of-a-long-distance-scale-across-ai-supercluster-302821351.html))_
- **Jul 6.** IDC Q1 2026 data: datacenter Ethernet switch revenue up 61% YoY to ~$10B against 3% server growth, with 800G sales up 10.3x to $3.58B — and NVIDIA now the top datacenter Ethernet vendor at $2.1B, ahead of Arista and Cisco _([The Next Platform (IDC); TechRepublic (Dell'Oro)](https://www.nextplatform.com/connect/2026/07/06/as-goes-ai-compute-so-goes-ethernet-networking/5267007))_
- **Jul 8.** Marvell published Keysight-validated Ultra Ethernet results (packet trimming, Auto Load Balancing, UET) on Teralynx switches — the most concrete public UET-on-silicon data point ahead of UEC-native NIC availability _([Marvell Blog](https://www.marvell.com/blogs/marvell-teralynx-packet-trimming-alb-uet-validation.html))_
- **Jul 9.** IEEE Spectrum: NVLink Fusion's photonics partners (Ayar Labs, Lightmatter, Marvell) signal optics moving into the scale-up domain as rack GPU density heads from 72 toward as many as 576 by 2027 _([IEEE Spectrum; Ayar Labs; Lightmatter](https://spectrum.ieee.org/nvlink-fusion-optics))_

**What this means.** After W27 strained the networking hypothesis, this week reinstated it with the strongest possible evidence: a commercial fabric product whose entire value proposition is monetizing the power constraint — stitching power-limited sites into one logical cluster (Metcalfe compounding Gilder). Architects planning 2027 capacity should now price scale-across fabric as a real alternative to waiting on single-site interconnection queues, and watch the 800G-to-1.6T ramp confirmed by 10.3x growth.

## Capital flow

| Category | Capital in | Revenue out | Burn:Revenue | Movement |
|---|---|---|---|---|
| Frontier Labs (OpenAI, Anthropic, Google DeepMind, xAI) | ~$95B (was ~$95B, flat) | ~$21B (was ~$21B, flat) | ~1.3x | No new primary capital; the froth moved to secondaries — Anthropic shares traded at an implied $1.2T (up 550% in a year, broker-reported), overtaking OpenAI's ~$908B — even as both labs kept cutting prices (Terra GA at half GPT-5.5; Grok 4.5 at $2/$6). |
| Hyperscaler-Hosted (Azure-OpenAI, AWS-Anthropic, Google Cloud-Gemini, Oracle-OCI) | ~$187B (was ~$187B, flat) | ~$62B (was ~$62B, flat) | ~3.0x | Meta committed C$13B (~$9.1B) to a 1 GW Alberta campus — its largest outside the US — fully funding its own 932 MW gas tolling deal because the grid cannot host multiple large AI loads; Microsoft cut 4,800 roles while funding its $2.5B Frontier AI-deployment unit. |
| Neoclouds (CoreWeave, Nscale, Crusoe, Lambda, Fluidstack, IREN) | ~$13.5B (was ~$13.5B, flat) | ~$5B (was ~$5B, flat) | ~2.7x | Delivery, not fundraising: Galaxy Digital completed Helios Phase I on schedule — 133 MW of critical IT load to CoreWeave under a 15-year lease with payments already flowing, against 526 MW committed across three phases and projected revenue above $1B a year. |
| On-Prem / Hybrid (Enterprise GPU clusters, sovereign and national programs, Cisco / Dell / HPE) | ~$94.5B (was ~$94B, up) | ~$36B (was ~$36B, flat) | ~2.6x | MARA acquired a 2 GW powered-land site in Matagorda County, Texas for up to $600M (8-K filed) — up to 1 GW of grid capacity by October 2027 — taking its potential portfolio to ~4.8 GW; Beijing pushed the Manus buyback toward a Tencent-led consortium at no less than $2B. |

### Frontier Labs — detail
The secondary prints are a sentiment signal, not a funding event: illiquid, scarce shares changing hands at ~5 buyers per 2 for Anthropic versus OpenAI. The strategic tension sharpened this week — valuations climbing toward IPO windows while the same labs reprice their core product downward and race to own the harness layer instead. Investors should discount secondary marks and watch what the labs ship (work runtimes, not just models) as the real pre-IPO positioning.

### Hyperscaler-Hosted — detail
The Alberta structure is the template to watch: hyperscalers becoming their own utilities rather than queueing for grid capacity, which drains the most credit-worthy anchor loads out of regulated interconnection processes. Meta's internal memo (via Reuters) targets 14 GW of compute in 2027 — double 2026 — with the Broadcom-designed Iris chip entering production in September. Buyers should read Microsoft's restructuring as the labor-side mirror of the same reallocation: sales headcount out, forward-deployed AI delivery capacity in.
**Transactions:**
  - **2026-07-08.** Meta: C$13B, 1 GW (expandable to 1.8 GW) Sturgeon County, Alberta data center with 932 MW Greenlight gas tolling agreement and 250 MW Capital Power ESA — ~$9.1B _([AP; Capital Power](https://nz.finance.yahoo.com/news/capital-power-enters-long-term-195300512.html))_

### Neoclouds — detail
The Helios delivery is the first clean data point that bitcoin-mine-to-AI conversion economics work on schedule and at contracted scale. After W27's Meta Compute repricing shock, the category needed evidence that booked capacity converts to revenue — this is it, though one delivery does not resolve the anchor-customer concentration question. Operators should keep diversifying; investors should watch CoreWeave's early-August Q2 print for backlog conversion velocity.
**Transactions:**
  - **2026-07-07.** Galaxy Digital Helios Phase I: 133 MW critical IT load delivered to CoreWeave, 15-year lease, lease payments began Q2 2026; Phase II 260 MW targeted H1 2027 — >$1B/yr projected _([FinanceFeeds](https://financefeeds.com/galaxy-helios-phase-i-133mw-coreweave/))_

### On-Prem / Hybrid — detail
Powered land is now a distinct asset class: MARA's $600M buys almost nothing but secured power rights, and the milestone-based structure prices energization risk explicitly. On the sovereignty side, the Manus unwind shows Beijing actively restructuring AI ownership — foreign capital out, domestic consortium in — while reportedly rationing H200 access below 200,000 units to favor domestic silicon. Enterprises with China exposure should treat model and vendor availability there as a policy variable, mirror-imaging the US release-gating regime.
**Transactions:**
  - **2026-07-09.** MARA Holdings: 2 GW powered-land site (1,200+ acres, Matagorda County, TX) from HIF USA; milestone payments, developed with Starwood Digital Ventures — up to $600M _([SEC 8-K; MARA press release](https://ir.mara.com/sec-filings/all-sec-filings/content/0000950142-26-002012/eh260804074_8k.htm))_

## Signal vs noise

- **Score 5/5 —** GPT-5.6 reached GA with Terra priced at $2.50/$15 per MTok — exactly half of GPT-5.5 — confirming the closed-lab repricing cycle.
  - _Sources:_ OpenAI GA announcement and pricing page; Vellum benchmark analysis
  - _Read:_ Prediction p55 resolves as a hit seven weeks early. The repricing cycle is now confirmed from three vendors (Sonnet 5, Terra, Grok 4.5) — buyers should renegotiate inference contracts this month, before Sonnet 5's intro pricing lapses Aug 31 anchors the new floor.
- **Score 4/5 —** Harness choice now swings agent cost more than model choice — over 2x per task at equal quality per Databricks, ~10x via tuned harness profiles per LangChain/NVIDIA.
  - _Sources:_ Databricks Engineering (merged-PR benchmark); LangChain and NVIDIA blogs
  - _Read:_ Two independent, methodologically serious studies published the same week, one on a real multi-million-line codebase. This is the strongest procurement-relevant finding of the month: benchmark the harness, not just the model, and treat per-token rate cards as a poor proxy for cost.
- **Score 3/5 —** NVIDIA's ~600 kW Kyber rack for Rubin Ultra has slipped to 2028 on PCB-midplane manufacturability.
  - _Sources:_ SemiAnalysis (paywalled report); NVIDIA public denial via CNBC/Wccftech
  - _Read:_ A credible specialist source against an explicit vendor denial — unresolvable this week. The facility-planning implication is real either way: anyone designing 2027-28 halls around 600 kW/rack and 800V DC should hold a 190-230 kW contingency. TSMC's Jul 16 earnings commentary on advanced packaging may arbitrate.
- **Score 2/5 —** Anthropic is now 'worth' $1.2 trillion, overtaking OpenAI.
  - _Sources:_ Business Insider via secondary outlets; broker-reported secondary-market prints
  - _Read:_ Secondary trades on illiquid, scarce shares are sentiment, not valuation — the primary mark remains the $965B Series H. The durable fact is directional demand (reportedly ~5 buyers per 2 for OpenAI); wait for the IPO range to treat any trillion-dollar figure as real.
- **Score 1/5 —** GPT-5.6 Sol Ultra proved the 50-year-old Cycle Double Cover Conjecture in under an hour.
  - _Sources:_ OpenAI proof PDF and prompt release; no independent verification
  - _Read:_ Noise until verified: no independent mathematical review, no Lean/Coq formalization, and the released prompt instructed the model to assume a proof exists — a setup that invites confident invalid arguments. If formal verification lands, this becomes the AI-research event of the quarter; until then, do not cite it.


## Levers

| Metric | Current | Prior | Direction | Threshold |
|---|---|---|---|---|
| Frontier lab cash position (avg months runway, top 3) | ~34-37 mo; Anthropic at implied $1.2T on secondaries (broker-reported), IPO calendars unchanged | ~34-37 mo; OpenAI leaning 2027 IPO, Anthropic holding Oct 2026 | flat | <18 mo triggers re-rating risk |
| Hyperscaler capex / AI revenue ratio (top 4 weighted) | ~5.0-5.3; Meta targets 14 GW of compute in 2027 (2x 2026) on $125-145B capex guidance | ~5.0-5.3; free-cash-flow crossover framed for ~Q3 2026 | flat | >6.0 invites investor pushback at next earnings |
| CoreWeave revenue backlog | ~$100B reported; Helios Phase I (133 MW) delivered on schedule — backlog now converting to lease revenue | ~$100B reported; Meta Compute repriced concentration risk (stock -12-15% Jul 1) | flat | Conversion velocity matters more than gross figure |
| NVIDIA Q-over-Q data center revenue | $75.2B Q1 FY27; Q2 guide $91B (reports Aug 26); SemiAnalysis sees H2 ~20% above consensus despite Kyber dispute | $75.2B Q1 FY27; Q2 guide $91B, reports Aug 26 | flat | Q2 FY27 guide $91B implies further +21% QoQ |
| Open vs closed gap on coding (SWE-Bench / agentic) | Effectively closed on cost-quality: GLM 5.2 statistically tied with Opus 4.8 at $1.28 vs $1.94/task (Databricks); Hy3 adds Apache-2.0 agentic-search lead | Narrowing from both sides: closed prices down (Sonnet 5 $2/$10; Terra promised at half GPT-5.5), open pressure sustained | down | Sustained open lead reshapes enterprise procurement |
| Sovereign AI commitments (count / aggregate $) | ~14 / ~$180B+ (flat; Meta Alberta and MARA Texas are corporate capital on power-rich land, not sovereign programs) | ~14 / ~$180B+ (flat; SB Neo's 10GW is corporate, not sovereign) | flat | — |
| PJM 2026/27 capacity auction price ($/MW-day) | $329.17; 2028/29 BRA bids closed Jul 7 — results post Jul 14 after 4 p.m. ET (prediction p54 resolves) | $329.17; 2028/29 BRA bids close Jul 7, results Jul 14 (slipped ~1 week) | flat | 11x in 24 months — power is the new binding constraint |
| Time-to-power, busiest US markets (months) | 60-84; hyperscalers routing around queues — Meta fully funds its own generation in Alberta because the grid cannot host multiple large loads | 60-84; FERC intervenor deadline Jul 9, tariff responses due Aug 17 | flat | — |
| Cost-per-task, frontier reasoning model | ~$0.06-$0.12 effective; GPT-5.6 GA tiering (Sol $5/$30 / Terra $2.50/$15 / Luna $1/$6), Grok 4.5 at $2/$6 — and harness choice swings per-task cost 2x+ | ~$0.08-$0.13 effective; Sonnet 5 $2/$10 intro, GPT-5.6 tiering Sol $5/$30 / Terra $2.50/$15 / Luna $1/$6 | down | — |
| Custom silicon share of incremental AI compute | ~34-37%; Meta's Iris enters production in September, Broadcom-Apple extended through 2031 (8-K), AWS raising Trainium 3 orders 20-30% | ~33-36%; silicon-IP layer forming (Oxmiq) to lower custom-ASIC entry cost | up | >35% materially compresses merchant GPU pricing |
**Lever detail:**
- **Frontier lab cash position (avg months runway, top 3).** Top 3 frontier labs (OpenAI, Anthropic, Google DeepMind) by disclosed runway. No new primary financing in-window; the movement was all in secondary marks — Anthropic at an implied $1.2T (up 550% YoY) versus its $965B Series H primary, with OpenAI around $908B on the same platforms. Both S-1s remain confidential. Boards should not read secondary froth as balance-sheet strength; the runway math is unchanged.
- **Hyperscaler capex / AI revenue ratio (top 4 weighted).** Top 4 hyperscalers (MSFT, GOOG, META, AMZN) weighted aggregate of capex divided by AI-attributable revenue. No top-4 print in-window, but Meta's internal memo (via Reuters) hardened the numerator: ~7 GW of compute in 2026 doubling to 14 GW in 2027, with the Alberta campus adding self-funded generation to the bill. Microsoft's 4,800-role cut alongside its $2.5B Frontier unit shows the offsetting cost discipline. The earnings wave starting the week of Jul 21 is the next test against actuals.
- **CoreWeave revenue backlog.** Booked but unrecognized revenue; official next print is Q2 in early August. This week supplied the first hard conversion data point since the Meta Compute shock: Galaxy delivered 133 MW of critical IT load on schedule under a 15-year lease (526 MW committed across three phases, >$1B/yr projected revenue). Conversion works; concentration risk remains — watch the Q2 print for anchor-customer mix.
- **NVIDIA Q-over-Q data center revenue.** No earnings event in-window, but two structural updates: SemiAnalysis reported the ~600 kW Kyber rack for Rubin Ultra slipping to 2028 (NVIDIA denied it the same day, holding to H2 2027), while separately projecting NVIDIA data-center compute revenue ~20% above consensus for H2 FY2027. Vera Rubin systems remain in full production shipping to eight cloud partners this fall. Demand is not the question; rack-level manufacturability is.
- **Open vs closed gap on coding (SWE-Bench / agentic).** The strongest week yet for this lever: Databricks' merged-PR benchmark put open GLM 5.2 in the top capability tier statistically tied with Opus 4.8 at two-thirds the per-task cost, on a real multi-million-line codebase with held-out test grading. Tencent's Hy3 (Apache 2.0, ~$0.20/$0.80 per MTok) leads open models on agentic search and went release-to-local-deployment in ~30 hours. The closed frontier still holds absolute SWE-Bench Pro leadership (Fable 5 at 80%+), but the cost-quality frontier is now open-weight territory.
- **Sovereign AI commitments (count / aggregate $).** Analyst-curated count of sovereign/national AI-compute commitments. No new drawn commitment in-window. The sovereignty story this week was restrictive rather than expansive: Beijing reportedly capping H200 purchases below 200,000 units to favor domestic silicon, and forcing the Manus ownership unwind toward a Tencent-led consortium at no less than $2B. Sovereign policy is shaping capital flows without deploying new capital.
- **PJM 2026/27 capacity auction price ($/MW-day).** The 2026/27 BRA cleared at the FERC cap ($329.17); 2027/28 at $333.44. The 2028/29 bid window closed on schedule Jul 7 with results due Jul 14 (cap ~$325, floor $175). PJM's proposed September-October backstop procurement against an anticipated shortfall still signals the operator itself planning for scarcity. Budget at-cap through 2028; a materially sub-cap print Monday would be the first crack in the pattern — and the week's highest-information event.
- **Time-to-power, busiest US markets (months).** Months from new-load interconnection request to energization. The headline structural development: Meta's Alberta campus pairs a 932 MW gas tolling agreement with a 250 MW Capital Power ESA because the grid explicitly cannot support multiple large AI loads — the self-funded-generation template that removes hyperscalers from the queue entirely. FERC tariff responses remain due Aug 17. For everyone without a generation balance sheet, the 60-84 month reality is unchanged; scale-across fabric (see networking lens) is emerging as the architectural workaround.
- **Cost-per-task, frontier reasoning model.** Median cost across frontier-tier reasoning models for a benchmark complex task. The ceiling dropped again (Terra GA at half GPT-5.5; Grok 4.5 claiming ~4.2x fewer output tokens per task), but the bigger update is structural: Databricks showed the same model at the same effort costs over 2x more per task depending on harness, and LangChain/NVIDIA hit near-Opus quality at ~$4.48 vs $43.48 per suite run via harness tuning alone. Route by measured cost-per-completed-task with the harness as an explicit variable.
- **Custom silicon share of incremental AI compute.** Three converging data points pushed this lever up: Meta's Broadcom-designed Iris MTIA chip enters production in September on a ~6-month cadence toward a 14 GW 2027 target; Broadcom filed an 8-K extending its Apple custom-ASIC partnership through 2031; and DigiTimes reports AWS telling suppliers to raise Q3 Trainium 3 shipments 20-30% above plan. The co-design duopoly (Broadcom, Marvell) keeps compounding — investors should note custom silicon is now the hyperscalers' stated path to doubling compute without doubling NVIDIA spend.

## Predictions

- **`p57-gemini-3-5-pro-ga-jul31` _[software]_ — Gemini 3.5 Pro reaches public general availability — a callable API model ID with published pricing — by July 31, 2026, after slipping past its June window and the reported July 17 target.**
  - Confidence: 58%. Deadline: By July 31, 2026.
  - Trigger: Google Gemini API model list / pricing page showing a GA gemini-3.5-pro model ID.
- **`p58-harness-cost-telemetry` _[software]_ — At least one major agent platform (OpenAI, Anthropic, GitHub, or Cursor) ships product-level per-task or per-harness cost telemetry or routing controls — beyond session budget caps — by August 31, 2026.**
  - Confidence: 64%. Deadline: By August 31, 2026.
  - Trigger: Product changelog or GA announcement exposing per-task cost measurement or harness-level cost controls.
- **`p59-tsmc-q2-capex-raise` _[hardware]_ — TSMC's July 16 Q2 earnings raise or reiterate the top end of full-year 2026 capex guidance and report HPC/AI platform revenue up more than 50% year over year, confirming the packaging-constrained AI capex ramp.**
  - Confidence: 62%. Deadline: By July 16, 2026.
  - Trigger: TSMC Q2 2026 earnings release and investor call (Jul 16; June revenue print Jul 13, typhoon-delayed).
- **`p60-scale-across-follow-on` _[networking]_ — A second named vendor or operator announces a commercial cross-data-center scale-across AI fabric deployment or product launch — following DriveNets/WhiteFiber — by September 30, 2026.**
  - Confidence: 61%. Deadline: By September 30, 2026.
  - Trigger: Vendor or operator press release for a commercial (not lab) multi-site training-fabric deployment; WhiteFiber's own Q3 commercial launch also qualifies if it lands with a named second customer.

### Prior predictions scored

- `p53-skhy-debut-validates-memory` _[capital]_ — **PARTIAL** — SK hynix's Nasdaq ADS offering prices at or above its indicated ~$166/ADS level and closes its first trading week above the offer price, by July 31, 2026. — Partial. The pricing clause missed — the offering priced at $149/ADS, below the ~$166 indication — but the market clause is on track: demand was ~7x available shares and day one closed up ~13% at ~$168, well above offer. The memory thesis was validated by the demand, not the price; regular trading (SKHY) begins Jul 13.
- `p54-pjm-2028-29-at-cap` _[power]_ — **PENDING** — The PJM 2028/29 base residual auction clears within 5% of the ~$325/MW-day cap when results post on July 14, 2026. — Pending — resolves Monday. The bid window closed on schedule Jul 7; results post Jul 14 after 4 p.m. ET. No in-window signal changed the at-cap setup, and PJM's proposed fall backstop procurement still implies the operator expects scarcity.
- `p55-terra-confirms-repricing-cycle` _[software]_ — **HIT** — GPT-5.6 reaches broad GA with the Terra tier priced at or below $2.50/$15 per MTok — half of GPT-5.5's rate — confirming a closed-lab repricing cycle rather than a one-off Sonnet 5 cut, by August 31, 2026. — Hit, seven weeks early. GPT-5.6 went GA Jul 9 with Terra at exactly $2.50/$15 per MTok. Grok 4.5's $2/$6 launch the day before makes it a three-vendor repricing cycle (Sonnet 5, Terra, Grok 4.5), not a one-off.
- `p56-samsung-hbm4-to-nvidia` _[hardware]_ — **HIT** — Samsung's HBM4 supply to NVIDIA is publicly confirmed — via earnings call, company statement, or multi-source supply-chain reporting — by August 31, 2026. — Hit on the multi-source-reporting trigger: Korean press (Seoul Economic Daily, Korea Herald) reported alongside Samsung's record Q2 guidance that HBM4 — in mass production since February for NVIDIA's Vera Rubin — reached $1B in sales within four months. Caveat: Samsung's Jul 30 divisional results would make it unambiguous from the company itself.

## Synthesis

### Connecting the dots

- **The enterprise procurement unit shifted from the model to the harness this week: OpenAI and Anthropic shipped competing general-work runtimes 48 hours apart, and two independent quantitative studies showed the harness now swings agent economics more than the model does.** _[abductive, 74% confidence]_
  1. Anthropic pushed Claude Cowork to web and mobile with cloud-run background sessions on Jul 7, citing 1.2M sessions across 600,000+ organizations showing most Cowork use is non-coding knowledge work.
  2. OpenAI answered on Jul 9 with ChatGPT Work — the Codex task runtime generalized to all knowledge work — bundled into a desktop app available on every plan including Free, explicitly citing 1M+ of Codex's 5M weekly users working outside software.
  3. Databricks' merged-PR benchmark found the same model at the same effort costs over 2x more per task in one harness than another at equal quality, and LangChain/NVIDIA showed harness tuning alone lifts an open model to near-Opus quality at roughly 10x lower cost.
  Evidence: [OpenAI ChatGPT Work launch](https://openai.com/index/chatgpt-for-your-most-ambitious-work/), [Anthropic Cowork web/mobile](https://claude.com/blog/cowork-web-mobile), [Databricks coding-agent benchmark](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase), [LangChain/NVIDIA harness-tuning playbook](https://www.langchain.com/blog/tuning-the-harness-not-the-model-a-nemotron-3-ultra-playbook)
- **The physical-scarcity trade W27 framed got marked to market with real money in a single week — public-market, corporate-treasury, and hyperscaler capital all paid up for memory and power at once, confirming that margin is pooling in scarce physical inputs while the model layer price-wars.** _[deductive, 78% confidence]_
  1. SK hynix closed the largest-ever foreign US IPO at $26.5B with demand reported at 7x available shares and a first-day close up ~13%, while Samsung guided to a record ~KRW 89.4T quarter on AI memory with HBM4 hitting $1B in sales within four months.
  2. Meta committed C$13B to a 1 GW Alberta campus and — because the grid cannot host multiple large AI loads — is funding its own 932 MW gas tolling deal plus a 250 MW supply agreement, while MARA paid up to $600M (8-K filed) for 2 GW of powered land in Texas.
  3. The same week, the model layer kept cutting prices: GPT-5.6 GA'd with Terra at $2.50/$15 (half of GPT-5.5), and Grok 4.5 launched at $2/$6 positioning explicitly on cost-per-task.
  Evidence: [SK hynix $26.5B IPO close](https://techcrunch.com/2026/07/10/sk-hynix-raises-26-5b-in-the-biggest-foreign-ipo-in-us-history-is-urged-to-build-new-us-fabs/), [Samsung record Q2 guidance](https://news.samsung.com/global/samsung-electronics-announces-earnings-guidance-for-second-quarter-2026), [Meta Alberta 1 GW with dedicated generation](https://nz.finance.yahoo.com/news/capital-power-enters-long-term-195300512.html), [MARA 2 GW powered-land 8-K](https://ir.mara.com/sec-filings/all-sec-filings/content/0000950142-26-002012/eh260804074_8k.htm)
- **The open-weight floor is rising through deployment economics, not just benchmarks: a permissively licensed near-frontier model went from release to local-hardware viability in about 30 hours, and enterprise-grade evidence now shows open models at frontier task quality for a third of the cost.** _[inductive, 71% confidence]_
  1. Tencent released Hy3 (295B MoE, 21B active) under a clean Apache 2.0 license on Jul 6; community GGUF quants with 1M context landed within ~30 hours, and a llama.cpp pull request using the model's MTP layer for speculative decoding measured +40% local throughput.
  2. Databricks' internal benchmark put open-weight GLM 5.2 statistically tied with Claude Opus 4.8 on quality at $1.28 vs $1.94 per task, and the colibri project demonstrated the 744B GLM-5.2 running in 25GB of consumer RAM by streaming experts from disk.
  3. Distribution caught up the same week: Hugging Face's curated open-weight collection landed in Microsoft Foundry with one-click managed-compute deployment, CVE-scanned runtimes, and per-deployment billing.
  Evidence: [Tencent Hy3 model card](https://huggingface.co/tencent/Hy3), [llama.cpp Hy3 MTP speculative decoding PR](https://github.com/ggml-org/llama.cpp/pull/25395), [Databricks GLM 5.2 cost-quality parity](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase), [Hugging Face collection in Microsoft Foundry](https://huggingface.co/blog/microsoft/foundry-managed-compute)
- **Frontier availability is now gated by governments on both sides of the Pacific: GPT-5.6 reached GA only after a government-coordinated preview with unpublished evaluation criteria, while Beijing simultaneously rationed NVIDIA H200 access and forced the Manus ownership unwind toward a Tencent-led consortium.** _[inductive, 66% confidence]_
  1. GPT-5.6 went GA on Jul 9 after a 12-day government-coordinated restricted preview; the system card discloses that Commerce's CAISI ran pre-deployment evaluations whose criteria remain unpublished, and all three tiers are rated High in bio/chem and cyber.
  2. The Information and Reuters reported Beijing preparing to let Alibaba, ByteDance, and DeepSeek buy H200s — but capped below 200,000 units, less than half of what was requested, after months of withheld approvals favoring domestic silicon.
  3. Tencent entered talks to lead a consortium buying Manus back from Meta at no less than $2B — the direct consequence of Beijing ordering Meta's acquisition unwound in April.
  Evidence: [GPT-5.6 system card and CAISI evaluations](https://deploymentsafety.openai.com/gpt-5-6), [China H200 quota reporting](https://www.thestandard.com.hk/finance/article/336792/China-plans-to-let-top-AI-firms-buy-limited-Nvidia-H200-chips-the-Information-reports), [Tencent-Manus buyback talks](https://the-decoder.com/tencent-moves-to-buy-majority-stake-in-manus-after-beijing-forced-meta-to-unwind-its-2-billion-deal/)

### Thesis test

- **Hypothesis 1 — The cycle is accelerating, not slowing.** — **SUPPORTED**. Four frontier-relevant releases landed in four days — GPT-5.6 GA (Jul 9), Grok 4.5 (Jul 8), Muse Spark 1.1 with Meta's first-ever model API (Jul 9), and Tencent's Apache-2.0 Hy3 (Jul 6) — and OpenAI set GPT-5.4's retirement for Jul 23, a deprecation cadence measured in months. The community-to-local pipeline compressed too: Hy3 went from release to quantized local deployment in ~30 hours, the fastest such cycle recorded for a 295B-class model. Cadence is compressing at both the frontier and the floor simultaneously. Evidence: [GPT-5.6 GA with GPT-5.4 retirement date](https://openai.com/index/gpt-5-6/), [Hy3 community GGUF quants ~30 hours after release](https://huggingface.co/satgeze/Hy3-1M-GGUF)
- **Hypothesis 2 — Capital is concentrated, returns are diffuse.** — **SUPPORTED**. The spread widened visibly this week: Anthropic traded at an implied $1.2T on secondary markets (up 550% in a year, broker-reported) while the labs kept cutting prices into that valuation — Terra at half GPT-5.5's rate, Grok 4.5 at $2/$6. Meanwhile the clearest realized returns landed at the physical layer (Samsung's record ~KRW 89.4T quarter, SK hynix's 7x-oversubscribed IPO) and in labor substitution (Microsoft cutting 4,800 roles while funding a $2.5B AI-deployment unit). Capital pools at the model layer; this week's cash profits showed up in memory, power, and restructured cost bases. Evidence: [Anthropic $1.2T implied secondary valuation](https://thenextweb.com/news/anthropic-1-2-trillion-secondary-valuation-openai), [Samsung record AI-memory quarter](https://news.samsung.com/global/samsung-electronics-announces-earnings-guidance-for-second-quarter-2026), [Microsoft cuts alongside $2.5B Frontier unit](https://www.cnbc.com/2026/07/06/microsoft-cuts-2point1percent-of-employees-as-xbox-unit-plans-to-spin-studios.html)
- **Hypothesis 3 — Networking is the durable layer.** — **SUPPORTED**. After W27 strained the hypothesis, this week delivered its strongest evidence in a month: DriveNets and WhiteFiber deployed the first commercial long-distance scale-across AI supercluster (111.2 Tbps over 83 km at 0.9 ms, within 8% of the physical limit of light in fiber), explicitly framed as an escape from single-site power constraints — networking directly monetizing the power bottleneck. The market data agrees: datacenter Ethernet switch revenue grew 61% YoY against 3% server growth, with 800G sales up 10.3x, and NVLink Fusion's photonics partners signal optics entering the scale-up domain as racks head toward 576 GPUs. Evidence: [DriveNets/WhiteFiber scale-across supercluster](https://www.prnewswire.com/il/news-releases/drivenets-announces-industrys-first-commercial-deployment-of-a-long-distance-scale-across-ai-supercluster-302821351.html), [IDC Q1 2026 Ethernet switch data](https://www.nextplatform.com/connect/2026/07/06/as-goes-ai-compute-so-goes-ethernet-networking/5267007)
- **Hypothesis 4 — Open weights pull the floor up.** — **SUPPORTED**. The mechanism operated end-to-end this week: Tencent shipped Hy3 under clean Apache 2.0 at ~$0.20/$0.80 per MTok, the community made it locally deployable within 30 hours (with +40% throughput from MTP speculative decoding), Databricks published enterprise evidence that open GLM 5.2 statistically ties Opus 4.8 at two-thirds the per-task cost, and Microsoft Foundry began one-click distribution of curated open weights into enterprise Azure estates. The floor is rising on quality, cost, deployability, and distribution simultaneously — and closed-lab pricing (Terra at half GPT-5.5, Grok 4.5 at $2/$6) is visibly responding. Evidence: [Hy3 Apache 2.0 release](https://huggingface.co/tencent/Hy3), [Databricks open-vs-closed per-task economics](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase)
- **Hypothesis 5 — Power is the binding constraint for the next 24 months.** — **SUPPORTED**. Meta's Alberta announcement is the cleanest single confirmation yet: the company is fully funding its own generation (a 932 MW gas tolling deal plus a 250 MW supply agreement) because the grid explicitly cannot support multiple large AI loads — the hyperscaler is becoming its own utility. MARA paid up to $600M for powered land whose value is entirely its 2 GW of secured capacity, Galaxy's 133 MW delivery to CoreWeave started a >$1B/year revenue stream, and the week's flagship networking deployment exists specifically to stitch power-constrained sites into one logical cluster. PJM's 2028/29 results (Jul 14) are the next hard test. Evidence: [Meta Alberta dedicated-generation structure](https://nz.finance.yahoo.com/news/capital-power-enters-long-term-195300512.html), [MARA powered-land acquisition 8-K](https://ir.mara.com/sec-filings/all-sec-filings/content/0000950142-26-002012/eh260804074_8k.htm)

### Pattern watch

- **Government action is a standing gate on frontier-model availability — and the gate is now bilateral.** _[inductive, 4 weeks observed]_
  - W25: Fable 5 and Mythos 5 suspended under US export controls (Jun 12).
  - W26: Mythos 5 restored only for ~100 'Annex A' critical-infrastructure organizations.
  - W27: GPT-5.6 previewed to ~20 government-vetted partners at the US government's request; Fable 5 restored globally after control withdrawal.
  - W28: GPT-5.6 GA'd only after a 12-day government-coordinated preview with CAISI pre-deployment evaluations (criteria unpublished); Beijing simultaneously rationed H200 purchases below 200,000 units and pushed the Manus ownership unwind.
  Next week: Gemini 3.5 Pro's reported Jul 17 GA is the test: if it ships without a government-coordinated preview phase, the US gate is OpenAI/Anthropic-specific rather than industry-standard; if it gets the same treatment, pre-release government review has become the de facto US frontier release process.
- **Agent economics are decoupling from model price lists — first tokens-per-task, now the harness itself.** _[inductive, 3 weeks observed]_
  - W26-W27: Artificial Analysis showed Sonnet 5's token appetite makes it cost more per completed task than the nominally pricier Opus 4.8, splitting sticker price from cost-per-task.
  - W27: Grok-class and GPT-5.6-tier pricing moves made cost-per-task the explicit competitive axis.
  - W28: Databricks showed the same model at the same effort costs over 2x more per task depending on harness choice, and LangChain/NVIDIA showed harness tuning alone closes most of the open-vs-closed quality gap at ~10x lower cost.
  Next week: Within two weeks another major platform (GitHub, Cursor, or a model vendor) publishes per-task or per-harness cost telemetry, or ships harness-level cost controls — confirming harness engineering as the new cost-optimization layer. If instead pricing discussion stays at per-token rate cards, the pattern is ahead of the market.
- **Physical-input scarcity (memory, power) keeps marking itself to market with progressively harder money.** _[inductive, 4 weeks observed]_
  - W25-W26: all three HBM makers volume-shipping HBM4; Micron confirmed 2026 supply fully contracted and HBM4 ramping ~2x faster than HBM3E.
  - W27: SK hynix filed a ~$29.4B Nasdaq listing and reportedly removed price caps from long-term memory contracts.
  - W28: the IPO closed at $26.5B with 7x demand and a +13% first-day pop; Samsung guided to a record ~KRW 89.4T quarter; Micron raised US investment plans to $250B+; Meta and MARA paid premiums for secured power.
  Next week: PJM's 2028/29 base residual auction results (Jul 14, after 4 p.m. ET) clear within 5% of the ~$325/MW-day cap, per prediction p54. A materially sub-cap print would be the first hard counter-evidence to the power-scarcity leg of this pattern.

### Second-order effects

- **Trigger:** OpenAI bundles ChatGPT Work (with Codex) into a desktop app available on every plan including Free, while Anthropic ships Cowork to web and mobile. **Effect:** Agent-harness distribution collapses into the subscription suites, squeezing standalone agent startups on distribution rather than capability — and enterprise governance becomes the real negotiation surface, since OpenAI's own rollout ships Work off-by-default with a two-week admin preview. Procurement teams that treated agents as a tool category must now treat them as a suite default that arrives enabled unless someone opts out. _(Horizon: Q4 2026. Who moves: Standalone agent-product startups, CIOs and IT admins managing suite defaults, and vertical-AI vendors whose wedge was 'the agent' rather than the workflow.)_
- **Trigger:** Databricks and LangChain/NVIDIA publish hard evidence that harness choice swings per-task cost 2-10x at equal quality. **Effect:** Enterprises re-run agent cost benchmarks and discover open-model-plus-tuned-harness parity, shifting spend from closed-model API contracts toward serving infrastructure and harness engineering as a discipline. Closed labs respond by bundling harness and model more tightly (exactly what ChatGPT Work does), making the harness a lock-in layer just as the model layer commoditizes. _(Horizon: H2 2026. Who moves: Enterprise AI platform teams, closed-lab API revenue, open-model serving providers, and anyone budgeting agent fleets on per-token rate cards.)_
- **Trigger:** Meta fully funds dedicated generation for its Alberta campus because the grid cannot host multiple large AI loads. **Effect:** Self-funded, behind-the-meter generation becomes the hyperscaler template for new campuses, which drains the most credit-worthy anchor loads out of utility interconnection queues — leaving regulated grid processes (and the FERC/PJM docket calendar) to govern everyone else. Non-hyperscale buyers inherit longer queues and higher capacity prices while hyperscalers effectively secede from the constraint. _(Horizon: 2027-2028. Who moves: Utilities and grid operators, colocation and enterprise data-center developers without generation balance sheets, and state energy regulators.)_

### Strategic outlook

This week hardened the 12-month posture on both ends of the stack. At the physical layer, the scarcity trade is no longer a thesis — it is a closed $26.5B IPO with 7x demand, a record memory quarter, and hyperscalers buying their own power plants; treat memory and secured power as strategic inventory, and expect the Jul 14 PJM print to confirm at-cap capacity pricing through 2028. At the work layer, the unit of enterprise AI procurement just shifted from the model to the harness: ChatGPT Work and Claude Cowork will land inside your organization through suite defaults, not RFPs, so governance capacity — admin opt-outs, budget caps, audit streaming — is the binding internal constraint to build now. And run the open-weight math again: with Apache-2.0 Hy3 deployable locally in 30 hours, GLM 5.2 tying Opus quality at two-thirds the per-task cost, and harness tuning worth more than model choice, the cost floor for capable agents is falling faster than closed-lab price cuts — which is precisely why the closed labs are racing to own the harness instead.



## Watchlist

- **Jul 13-16 — TSMC June revenue (Jul 13, typhoon-delayed) and Q2 earnings (Jul 16).** The first hard AI-capex read of the season — capex guidance, HPC platform growth, and advanced-packaging commentary will also arbitrate the SemiAnalysis-vs-NVIDIA Kyber dispute (prediction p59).
- **Jul 14 — PJM 2028/29 capacity auction results (after 4 p.m. ET).** Prediction p54 resolves: an at-cap clearing confirms power as the binding constraint through 2028; a materially sub-cap print would be the first crack in the pattern and would reprice siting strategy.
- **Jul 17 — Gemini 3.5 Pro reported GA target.** Two slips already (June, then Jul 17); prediction p57 tracks GA by Jul 31. Also the test of the government-preview pattern: if Google GAs without a CAISI-coordinated phase, pre-release review is an OpenAI/Anthropic-specific regime, not an industry standard.
- **Jul 23-24 — GPT-5.4 retirement (Jul 23) and DeepSeek legacy alias shutdown (Jul 24).** Two hard migration deadlines a day apart — the closed and open ecosystems now deprecate at the same aggressive cadence, and both will surface integration debt in agent fleets built on pinned model IDs.
- **Week of Jul 21 — Q2 earnings wave opens: ServiceNow (Jul 22), then Alphabet / Microsoft / SAP (dates aggregator-estimated).** First top-4 hyperscaler prints of the season test the free-cash-flow-crossover narrative against actuals; Microsoft's report is the one to watch for Copilot revenue disclosure and Frontier-unit framing after the 4,800-role cut.
- **Jul 30 — Samsung Q2 divisional results.** The company-level HBM4 disclosure (reported $1B in four months, ~$10B annualized pace by year-end) would convert prediction p56's hit from supply-chain reporting to primary confirmation — and set the tone for SK hynix's first earnings as a US-listed company.

## Changelog

- W28 adds the harness-as-product story (ChatGPT Work vs Claude Cowork, Databricks and LangChain/NVIDIA harness economics), the SK hynix IPO close, Samsung's record memory quarter, and the Meta Alberta self-funded-generation template; p55 and p56 scored as hits, p53 partial, p54 resolves Jul 14.
- Research pipeline change: community discovery channels (r/LocalLLM, r/LocalLLaMA, Hacker News) added as monitored sources this week — they surfaced the Hy3 local-deployment pipeline, the Databricks harness benchmark, and the colibri GLM-5.2 project ahead of mainstream coverage, each verified against primary sources before grading.

---

Source of truth: `src/data/industry/weekly/2026-W28.ts`. Canonical HTML: <https://brianletort.ai/industry/weekly/2026-W28>. PDF: <https://brianletort.ai/downloads/ai-stack-weekly-2026-W28.pdf>.
