Skip to content

Cross-stack flywheel

AI Stack Weekly

For officers tracking AI market movement.

Model pricing is converging from both directions onto one frontier band — closed labs cutting down into $3/$15 while Chinese open-weight labs raise up into it — and the differentiation is shifting from price to who actually ships weights.

Abstract editorial illustration: three interlocking rings of cyan light — software, silicon, and network — turning as one flywheel on a dark field.
Issue 13 — this week's arc ran through all three lenses.

Executive summary

7 minute read

Key takeaways

  • Frontier model pricing converged from both directions onto one $1-3 input / $6-15 output band — closed labs cut down into it last week, Chinese open-weight labs raised up into it this week. Differentiation shifts from price to who actually ships weights.
  • Kimi K3 became the first open-weights-slated model to reach the closed-frontier tier on independent evaluations (57.1 on the AA Index, #4 of 189) — but launched hosted-only, with weights and license promised by July 27.
  • Compute began trading across former battle lines: Anthropic in reported talks to lease up to ~$10B of capacity from Meta, while GPU-backed credit institutionalized at rated spreads (Nebius $775M, Fitch-rated CoreWeave $2.6B DDTL).
  • The binding constraint migrated downstream of silicon and capital: TSMC's CEO said packaging is 'limiting my customers' growth', PJM cleared at the FERC cap 6,831 MW short, and New York signed the first statewide data-center permit moratorium.
  • House measurement: PJM's price cap is masking ~$11.6B/yr of capacity-market scarcity cost — about $84,000 per MW-year that AI-load budgets built on the capped price never see.
  • Watch July 27: the Kimi K3 weights-and-license drop is the single highest-information event in view. If it lands clean, frontier-class capability becomes self-hostable for the first time.

By the numbers

Kimi K3 per MTok — the flagship price band both sides now share
$3 / $15 — Roughly 3x its predecessor and at Claude Sonnet 5's standard rate card
K3 on the AA Intelligence Index — #4 of 189, ~3 pts off the closed frontier
57.1 — First open-weights-slated model at the closed-frontier tier
TSMC record Q2 revenue, +33.7% YoY; FY26 capex raised to $60-64B
$40.2B — SEC-filed; advanced packaging named as the binding constraint
PJM's shortfall as the auction cleared at the FERC cap a third straight year
6,831 MW — Only 525 MW of new generation cleared
Scarcity cost PJM's price cap masks — this issue's house measurement
~$11.6B/yr — ~$84,000 per MW-year of suppressed capacity cost
Meta's Hyperion expansion — the largest disclosed single-site AI investment
>$50B — 5 GW in Louisiana, up from $27B / 2 GW

Big story

Here is the claim nobody else made this week: frontier model pricing is now converging from both directions onto a single $1-3 input / $6-15 output band. Last week the closed labs cut down into it (GPT-5.6 Terra at half GPT-5.5's rate, Grok 4.5 at $2/$6). This week the open-weight side raised up into it — Moonshot priced Kimi K3 at $0.30/$3/$15 per MTok, roughly 3x its predecessor and at Claude Sonnet 5's standard rate card (though Sonnet's $2/$10 intro price runs through Aug 31, leaving K3 ~50% above the live street price) — while DeepSeek added the first peak-hour 2x surge pricing from a major model API. The 'end of cheap Chinese AI' framing itself was already circulating on launch day (the-decoder; Simon Willison called K3 the most expensive Chinese-lab model to date); what the convergence adds is the falsifiable market-structure read: with price no longer the differentiator, the axis shifts to deployment control — and K3 launched hosted-only with a vendor-reported 2.8T parameters, an unpublished license, and weights promised by July 27, the same week Thinking Machines' 975B Inkling shipped Apache 2.0 weights on day one. 'Open' is becoming a license-and-weights fact, not a launch label.

The capability side makes the pricing story matter: K3 is the first open-weights-slated model to reach the closed-frontier tier on independent evaluations — 57.1 on the Artificial Analysis Intelligence Index (#4 of 189, within three points of Claude Fable 5 and GPT-5.6 Sol) and #1 on LMArena's Frontend Code Arena. Meanwhile the physical layer tightened on schedule: TSMC posted a record $40.2B quarter, raised 2026 capex to $60-64B, added $100B to Arizona — and CEO C.C. Wei said advanced packaging capacity is 'so tight that now it's limiting my customers' growth.' PJM's 2028/29 capacity auction cleared at the FERC cap for the third straight year while 6,831 MW short (our house measurement below prices what the cap hides), and New York signed the first statewide data-center permit moratorium a day later. Compute itself began trading across former battle lines: Anthropic in reported talks to lease up to $10B of capacity from Meta, while GPU-backed credit institutionalized — a $775M Nebius facility at SOFR+2.50% (2.5 points over the floating benchmark rate), a Fitch-rated $2.6B CoreWeave delayed-draw term loan, and the first inference-ASIC-collateralized facility.

What to do with it: boards should re-price single-vendor model dependency against a credible self-host option arriving July 27 — and treat the pricing convergence as the negotiating window. Investors should note credit and equity are telling different stories about GPU assets (cleanly at CoreWeave; the sector read is contested — Nebius equity rose on its debt deal). Architects and operators should treat packaging, power, and permits — not chips or capital — as the binding constraints for 2027 planning.

Flywheel arc · all-three

'Open' is becoming a license-and-weights fact, not a launch label.

  • Frontier pricing is converging from both directions onto a single $1-3 input / $6-15 output band: closed labs cut down into it last week; this week Moonshot priced Kimi K3 at roughly 3x its predecessor and DeepSeek added the first peak-hour surge pricing from a major model API.
  • With price no longer the differentiator, the axis shifts to deployment control: K3 launched hosted-only with an unpublished license and weights promised by July 27 — the same week Thinking Machines' 975B Inkling shipped Apache 2.0 weights on day one.
  • The capability side makes the pricing story matter: K3 is the first open-weights-slated model to reach the closed-frontier tier on independent evaluations — 57.1 on the AA Intelligence Index (#4 of 189) and #1 on LMArena's Frontend Code Arena.
  • The physical layer tightened on schedule: TSMC's record $40.2B quarter came with an on-the-record packaging-capacity warning, PJM cleared at the FERC cap 6,831 MW short, and New York signed the first statewide data-center permit moratorium a day later.
  • What to do: treat the pricing convergence as the negotiating window on single-vendor model dependency, and treat packaging, power, and permits — not chips or capital — as the binding constraints for 2027 planning.

Software lens

What this means

The Jevons arc ran hot this week: near-frontier capability got dramatically cheaper to acquire (K3 independently measured at $0.94 per Index task; Inkling free to self-host under Apache 2.0), and both flagship open releases ship 4-bit-native targeting very large self-hosted footprints — Moonshot recommends 64+ accelerator supernodes, Inkling needs ~600GB in NVFP4 — pulling inference demand toward private clusters and their memory and interconnect fabrics. Architects should hold procurement decisions on the K3 tier until the July 27 weights-and-license drop resolves, and see the Model Pulse for the full architecture read on why routing stability and prompt-cache hit rates are the new cost levers.

  • Near-frontier capability got dramatically cheaper to acquire: K3 independently measured at $0.94 per Index task; Inkling free to self-host under Apache 2.0.
  • Both flagship open releases ship 4-bit-native targeting very large self-hosted footprints — pulling inference demand toward private clusters and their memory and interconnect fabrics.
  • Hold procurement decisions on the K3 tier until the July 27 weights-and-license drop resolves.

Jul 13

OpenAI's GPT-5.6 family (Sol, Terra, Luna) reached GA on Amazon Bedrock via the bedrock-mantle Responses API at OpenAI first-party rates counting toward AWS commitments — with 272K context, 90% prompt-cache discounts, and Sol limited to two US East regions

Sources AWS What's New; AWS Bedrock model cards

Jul 15

Thinking Machines Lab released Inkling — a 975B-parameter (41B active) Apache 2.0 multimodal MoE with 1M context, the largest US-origin open-weights model — with weights live on Hugging Face day one, BF16 plus calibrated NVFP4 checkpoints, and day-0 vLLM/SGLang/llama.cpp support

Sources Hugging Face blog (co-published); Latent Space

Jul 15

OpenAI disclosed GPT-Red, an internal-only self-play RL red-teaming model that beat human red-teamers 84% vs 13% on indirect prompt-injection scenarios and hardened GPT-5.6 Sol to a 0.05% failure rate on its own attack corpus (all figures vendor-reported; preprint promised) — and said it will not be released due to offensive capability

Sources OpenAI; The New Stack; Help Net Security

Jul 16

Moonshot AI launched Kimi K3 — a vendor-reported 2.8T-parameter MoE (16 of 896 experts active) with 1M context, native text/image/video input, and Kimi Delta Attention — hosted-only at $0.30/$3/$15 per MTok, with open weights promised by July 27

Sources Moonshot/Kimi blog; Simon Willison; MarkTechPost

Jul 16

Kimi K3 debuted at 57.1 on the Artificial Analysis Intelligence Index (#4 of 189, vs Fable 5's 59.9 and GPT-5.6 Sol's 58.9) and took #1 on LMArena's Frontend Code Arena at 1,679 — the first open-weight-lab model to top an LMArena flagship board

Sources Artificial Analysis; Arena.ai leaderboard

Hardware lens

What this means

Both in-window prints confirmed the same structure: demand visibility is lengthening (TSMC guiding to slightly above 40% full-year growth; ASML nearly fully booked on 2027 EUV) while the binding constraint migrated downstream to advanced packaging — the CEO of the world's foundry said so on the record. Operators should treat Wei's packaging warning as the arbiter of the Rubin rack-delivery dispute, and watch Intel's High-NA production first as the first credible fork in leading-edge lithography cadence in a decade; if it converts to 14A customer wins, the Huang's-law slope gets a second supplier.

  • Demand visibility is lengthening: TSMC guiding to slightly above 40% full-year growth; ASML nearly fully booked on 2027 EUV.
  • The binding constraint migrated downstream to advanced packaging — the CEO of the world's foundry said so on the record.
  • Intel's High-NA production first is the first credible fork in leading-edge lithography cadence in a decade; 14A customer wins would give the Huang's-law slope a second supplier.

Jul 15

ASML Q2: EUR 9.3B net sales above the high end of guidance, system sales split 51% logic / 49% memory, ~65 Low-NA EUV shipments planned for 2026 (+45% YoY EUV system sales), 2027 Low-NA capacity 'close to fully covered' with +30% expansion planned

Sources ASML Q2 2026 press release and investor call

Jul 15

Intel Foundry became the first to ship high-volume logic patterned with High-NA EUV — Panther Lake layers on Intel 18A, dual-qualified in Oregon at yields matching the standard NXE platform; TSMC has said it won't run High-NA in production before ~2029

Sources ASML press release (joint with Intel Foundry)

Jul 15

NVIDIA's Huang, in Tokyo, denied the Rubin delay reports: 'Vera Rubin is already in production. Giant amounts of production incoming' — but named no customer-delivery date; Japan's FRONTia AI factory (27,500 Rubin GPUs) was announced the next day

Sources Tom's Hardware; multiple outlets

Jul 16

TSMC Q2: record $40.2B revenue (+33.7% YoY), HPC at 66% of revenue, 2nm at 3% of wafer revenue in its first material quarter, FY26 capex raised to $60-64B, $100B added to Arizona — and CEO C.C. Wei: advanced packaging capacity is 'so tight that now it's limiting my customers' growth'

Sources TSMC 2Q26 management report (SEC 6-K); earnings call

Jul 15

The Information: Google reportedly pitching TPUs to NVIDIA-centric neoclouds — offering to financially backstop TPU data centers and rent capacity back; Nebius, Lambda, and CoreWeave publicly demurred

Sources The Information via WinBuzzer (secondary)

Networking lens

What this means

The center of gravity this week was optics manufacturing capacity and deployment velocity rather than protocol wars: a $3B state-co-funded foundry bet on photonic ICs (Gilder — bandwidth supply industrializing ahead of demand) and a hyperscaler attacking the unglamorous fiber-connector bottleneck that throttles how fast bandwidth physically installs. A sovereign program standardizing on a single vendor's Ethernet fabric shows national fabrics becoming Metcalfe machines — every domestic lab and manufacturer that plugs in multiplies shared-fabric value, and deepens the lock-in debate at state scale. Watch the colo interconnect Q2 prints starting Jul 29 for the revenue confirmation.

  • The week's center of gravity was optics manufacturing capacity and deployment velocity, not protocol wars: a $3B state-co-funded photonics bet plus a hyperscaler attack on the fiber-connector bottleneck.
  • A sovereign program standardizing on a single vendor's Ethernet fabric shows national fabrics becoming Metcalfe machines — and deepens the lock-in debate at state scale.
  • Watch the colo interconnect Q2 prints starting Jul 29 for the revenue confirmation.

Jul 14

Tower Semiconductor announced a ~$3B (net of $1B Japan government grants) dual-track 300mm silicon photonics and SiGe capacity expansion in Japan with METI support — targeting $3.6B revenue in 2028 on optical interconnect demand

Sources Tower Semiconductor press release

Jul 15

3M and Microsoft announced a strategic partnership making Azure the first announced hyperscale cloud to deploy 3M Expanded Beam Optical fiber connectivity across AI data centers — attacking contamination-sensitive connector install and service time

Sources Microsoft Source newsroom; 3M investor release

Jul 16

Japan's METI-backed FRONTia project launched national AI infrastructure — a 140 MW Vera Rubin AI factory (27,500 Rubin GPUs, 13,750 Vera CPUs) scaled with Spectrum-X Ethernet fabric, with ¥1T (~$6.2B) of government support over five years

Sources NVIDIA Newsroom; Japan Times; NHK

Jul 16

Chelsio launched its seventh-generation AI Interconnect Platform with native 400Gb Ethernet and unified RDMA (RoCEv2 + iWARP) — widening the pool of standards-based NIC endpoints that can terminate open Ethernet AI fabrics

Sources Chelsio Communications press release

Capital flow

Capital in, revenue out, and the direction of travel.
CategoryCapital inRevenue outBurn to revenueMovement
Frontier Labs — OpenAI, Anthropic, Google DeepMind, xAI~$95B · prior ~$95B · flat~$21B · prior ~$21B · flat~4.5xNo new primary capital closed in-window; the action moved to compute sourcing — Anthropic in reported talks to lease up to ~$10B of capacity from Meta over two years (figures explicitly disputed as 'speculative' by one source), on top of its earlier $45B/3yr Colossus 1 access commitment.
Hyperscaler-Hosted — Azure-OpenAI, AWS-Anthropic, Google Cloud-Gemini, Oracle-OCI~$210B · prior ~$187B · up~$62B · prior ~$62B · flat~3.4xMeta expanded Hyperion (Louisiana) to 5 GW and >$50B — the largest disclosed single-site AI investment — while Google was confirmed as the customer behind the 2.7 GW 'Project Tembo' campus near Cheyenne, Wyoming (scalable to 10 GW); Meta simultaneously negotiated to sell capacity to Anthropic.
Neoclouds — CoreWeave, Nscale, Crusoe, Lambda, Fluidstack, IREN~$17B · prior ~$13.5B · up~$5B · prior ~$5B · flat~3.4xThe debt market institutionalized around AI compute in one week: Nebius closed a $775M GPU/contract-backed facility at SOFR+2.50% — 2.5 points over the floating benchmark rate, disclosed in an SEC 6-K (the filing form for foreign issuers) — Fitch rated CoreWeave's proposed $2.6B delayed-draw term loan (DDTL: credit committed now, drawn as GPUs are bought) at BB+ against take-or-pay contracts (customers pay whether or not they use the capacity), and General Compute secured the first inference-ASIC-collateralized facility (up to $400M against SambaNova chips).
On-Prem / Hybrid — Enterprise GPU clusters, sovereign and national programs, Cisco / Dell / HPE~$101B · prior ~$94.5B · up~$36B · prior ~$36B · flat~2.8xJapan launched FRONTia, the first national physical-AI infrastructure program — ¥1T (~$6.2B) of government support over five years for a 140 MW / 27,500-Rubin-GPU AI factory built by a 48-company consortium — while New York moved the other direction with the first statewide moratorium on discretionary environmental permits for data centers ≥50 MW.

Frontier Labs detail

If a top-2 lab will rent training and inference capacity from a rival's fleet, the labs have become compute-supplier-agnostic — compute is turning into a traded commodity across former battle lines, which weakens hyperscaler exclusivity moats and validates the neocloud business model even as it threatens neocloud pricing. Investors should treat the $10B figure as sentiment, not capital flow (no term sheet exists), but treat the direction — labs renting anyone's GPUs, Meta seeking external compute revenue — as real and structural.

Capital in value
$95B
Revenue out value
$21B
  • 2026-07-17 · Anthropic-Meta compute lease talks reported: up to ~$10B over two years, monthly payments, either side can walk; no term sheet — figures called 'speculative' by one source · ~$10B reported (disputed)

Sources NYT; CNBC; CNN

Hyperscaler-Hosted detail

Meta's twin moves — doubling Hyperion to >$50B on Monday while negotiating to lease capacity to Anthropic on Friday — mark the week hyperscaler capex stopped being purely internal: a fifth cloud is being born out of surplus AI capex (the hyperscaler-turned-compute-vendor role reversal circulated with the NYT's Jul 17 report; the pricing read is ours), and it prices against neoclouds, not Azure/AWS retail. The unresolved question is financing: no JV partner has been named for the Hyperion expansion, and if Meta funds it on-balance-sheet the depreciation drag lands directly on earnings. Buyers should watch the Jul 29 Microsoft print and late-July Alphabet/Meta calls for whether capex guides absorb these adds.

Capital in value
$210B
Revenue out value
$62B
  • 2026-07-13 · Meta: Hyperion (Richland Parish, LA) expanded to 5 GW and >$50B ex-chips, up from $27B/2 GW · >$50B
  • 2026-07-14 · Google confirmed as customer behind 2.7 GW 'Project Tembo' near Cheyenne, WY (via Jupiter Star Holdings; scalable to 10 GW), per county planning documents · ~$50B reported

Sources CNBC; SiliconANGLE (Meta post) · DCD; Cowboy State Daily

Neoclouds detail

At CoreWeave, credit and equity are telling opposite stories about the same assets: a rated take-or-pay DDTL signals lender confidence in contracted compute cash flows at the exact moment public equity reprices its concentration risk downward (stock down ~35% since the Meta Compute report — secondary-press figure). The honest caveat: this is not yet a sector-wide divergence — Nebius equity rose ~8% on its own debt announcement, with the facility read as easing the equity story. Investors should watch which story wins at CoreWeave's early-August Q2 print; operators should note that GPU-backed credit at institutional spreads materially lowers the capital cost of contracted capacity — for those with bankable offtake.

Capital in value
$17B
Revenue out value
$5B
  • 2026-07-17 · Nebius: first senior secured debt facility, GPU/contract-backed, SOFR+2.50%, matures Oct 2030, MUFG-led, oversubscribed (6-K filed; signed Jul 10) · ~$775M
  • 2026-07-16 · Fitch assigned CoreWeave's proposed DDTL 5.5 (SPV, GPU purchases against take-or-pay contracts) BB+/RR2; IDR affirmed BB- with positive outlook; backlog restated at $99.4B · $2.6B facility
  • 2026-07-17 · General Compute: debt facility from Upper90 collateralized by SambaNova inference ASICs — the first non-NVIDIA-GPU compute-backed facility · up to $400M ($100M initial)
  • 2026-07-15 · Crusoe + Lancium announced a 1 GW grid-connected AI campus in Childress, TX; construction from Q3 2026

Sources Nebius newsroom; Form 6-K · Fitch Ratings · TechCrunch; SiliconANGLE · DCD

On-Prem / Hybrid detail

FRONTia is small next to hyperscaler capex, but it sets the template other industrial states will copy: consortium plus national data plus a single-vendor reference design, explicitly scoped to physical AI and robotics. New York's EO 62 is the counter-template — narrower than the headlines suggest (discretionary state environmental permits only, completed applications exempt), but it hands every data-center opposition movement a governor-signed precedent one day after PJM confirmed the shortage that fuels the opposition. Enterprises siting on-prem capacity should now price state-policy risk alongside grid-queue risk.

Capital in value
$101B
Revenue out value
$36B
  • 2026-07-16 · Japan FRONTia: ¥1T (~$6.2B) government support over 5 years (¥387.3B FY2026); Noetra Corp. consortium (Sony, SoftBank, NEC, Honda + ~44 firms) + NVIDIA 140 MW Vera Rubin AI factory, online June 2028 · ~$6.2B / 5 yrs
  • 2026-07-14 · New York EO 62: first statewide moratorium on discretionary environmental permits for data centers ≥50 MW while the state develops a generic environmental impact statement

Sources NVIDIA/GlobeNewswire; Japan Times; NHK · NY Governor's office (EO text); Data Center Knowledge

See the full capital-flow breakdown

Signal vs noise

Signal score 5/5

PJM's 2028/29 capacity auction cleared at the $325/MW-day FERC cap for the third straight year — and still came up 6,831 MW short of the reliability requirement, with only 525 MW of new generation clearing.

Real, and the headline understates it: the price fell 2.5% only because the cap itself fell, while PJM's own uncapped simulation cleared 71% higher ($554.72 RTO / $776.69 ComEd) — the gap between capped and simulated price has widened three auctions running (a trendline Modo Energy charted first; our house measurement dollarizes it). Anyone budgeting AI load in the largest US grid should watch PJM's FERC filings this month for the backstop auction and data-center connect-and-manage framework, not the headline price.

Sources
PJM 2028/2029 BRA results report (primary); Utility Dive; Modo Energy

Signal score 4/5

TSMC will spend another $100B in Arizona (total $265B) after a record Q2 — revenue $40.2B (+33.7% YoY), net income +77% YoY, FY26 capex raised to $60-64B.

The financial capacity is audited and extraordinary; the commitment is real but the timeline is elastic — CEO C.C. Wei explicitly tied fab timing to 'how market conditions develop,' and the announcement's political utility is part of its purpose. Treat $265B as a ceiling with option value, not a schedule; groundbreaking dates and tool orders are the confirmation to watch.

Sources
TSMC SEC 6-K and earnings call (Q2 results grade 5); NYT, Reuters (Arizona commitment)

Signal score 4/5

Kimi K3 reached the closed-frontier tier on independent evaluations — 57.1 on the Artificial Analysis Index (#4 of 189) and #1 on LMArena's Frontend Code Arena — as an open-weights model.

The placements are genuinely independent and real; the 'open' framing is promissory — K3 launched hosted-only, the license text is unpublished, the HF repo 404'd, and Artificial Analysis classifies it proprietary until the July 27 weight drop lands. Two pricing caveats: 'Sonnet parity' holds only against the standard rate card (Sonnet 5's $2/$10 intro runs through Aug 31, putting K3 ~50% above the live street price), and effective cost exceeds sticker — always-on max-effort thinking consumed 130M tokens on the Index eval versus a 63M peer average. Hold the procurement conclusion two weeks.

Sources
Artificial Analysis (independent eval); Arena.ai leaderboard; Moonshot launch blog

Signal score 3/5

New York became the first state to impose a data-center moratorium, halting AI buildout in the state.

Directionally real, rhetorically inflated — and the context most coverage missed cuts the other way: the order landed about a month after the legislature passed a stricter 20 MW moratorium bill the governor is expected to veto, making EO 62 arguably a softening maneuver against the legislature rather than an opposition victory. It pauses only discretionary state environmental permits for facilities ≥50 MW, exempting completed applications, ministerial permits, and local approvals. The significance is precedential — copycat orders in Virginia, Texas, or Georgia would change the read from symbolic to structural.

Sources
EO 62 text (grade 5 for the event); Data Center Knowledge; law-firm client alerts

Signal score 2/5

Meta and Anthropic have a $10B, two-year compute leasing deal.

Hype as priced: the talks are real (multiple outlets independently confirm talks), but the tradable 'fact' the market moved on — $10B — is explicitly disputed, terms are in flux, and no term sheet exists. A buyer-proposed monthly-payment lease with walk-away rights is an overflow-capacity option, not a strategic commitment. The direction it signals (labs renting rivals' GPUs, Meta seeking compute revenue) is real; the number is not yet.

Sources
NYT (three unnamed sources); CNBC; CNN — CNN's source calls the figures 'speculative'

House measurement

Filing-Derived

PJM's price cap is masking roughly $11.6B per year of capacity-market scarcity cost — about $84,000 per MW-year that AI-load budgets built on the capped price never see.

Method: House computation from PJM's 2028/29 Base Residual Auction results report (published Jul 14). PJM cleared 138,318 MW at the $325.00/MW-day administrative cap and disclosed that its own uncapped market simulation would have cleared at $554.72/MW-day RTO-wide. Suppressed spread: 554.72 - 325.00 = $229.72/MW-day. Dollarized, holding cleared quantity constant: $229.72 x 138,318 MW x 365 days = $11.60B/year. Sanity check: the same arithmetic at the capped price reproduces PJM's own disclosed $16.4B total auction cost, validating the method.

Implication: Anyone underwriting AI load in PJM territory off the headline $325 clearing price is budgeting to an administratively suppressed number, not the market's own scarcity signal. Treat ~$84k/MW-yr as latent capacity-cost exposure — it surfaces through the September backstop procurement, the pending data-center connect-and-manage FERC filings, bilateral capacity contracts, or a lifted collar. Boards comparing PJM sites against ERCOT or self-generation should run the comparison at the simulated price, not the capped one.

Caveats: PJM's uncapped figure is the operator's own counterfactual simulation, not an observed market outcome; holding cleared quantity constant is an approximation; and how much of a capacity charge reaches a specific data-center load depends on its LSE pass-through and zone. The number measures the size of the administrative suppression, not a bill anyone receives today.

Suppressed spread
$229.72/MW-day — PJM's uncapped simulation ($554.72) minus the capped clearing price ($325.00) — a 71% gap, the widest of the last three auctions (18% -> 59% -> 71%)
Annualized suppression, RTO-wide
~$11.6B/yr — vs the $16.4B the auction actually cleared at — the market signaled ~71% more scarcity cost than the cap allowed to print
Hidden premium per MW of load
~$84,000/MW-yr — the per-MW-year capacity cost the cap suppresses; obligations are UCAP-adjusted (unforced capacity — capacity discounted for expected outages), and constrained zones run higher (ComEd simulated at $776.69/MW-day)
Illustrative 100 MW AI load
$11.9M vs $20.2M/yr — annual capacity cost at the capped price vs at PJM's own uncapped simulation — an $8.4M/yr gap per 100 MW that current budgets do not carry

Sources PJM 2028/2029 Base Residual Auction results report · PJM press release (procurement totals)

Synthesis · Connecting the dots

Inductive · 74% confidence

With frontier pricing converging from both directions onto one band, price stops differentiating Chinese open-weight labs from closed vendors — the competitive axis shifts to deployment control, and 'open' becomes a license-and-weights fact rather than a launch label. Resolution point: DeepSeek V4's GA pricing and the Jul 27 K3 weights drop.

Steel-man: Two data points make a segmentation, not an era: Moonshot kept its K2.x value tiers on sale ($0.60-$0.95 input), DeepSeek's off-peak rates are unchanged, and one flagship repricing plus one surcharge could be premium-tier positioning rather than the end of the discount strategy. The claim survives narrowed: the FLAGSHIP price band has converged — but if DeepSeek V4's GA pricing resets to pre-K3 ultra-cheap levels, the convergence read fails and the price war resumes (that is prediction p66).

  • Moonshot priced Kimi K3 at $0.30/$3/$15 per MTok — roughly 3x its K2.6 flagship tier and matching Claude Sonnet 5's standard rate card (though ~50% above Sonnet's live $2/$10 intro price, which runs through Aug 31) — the same day independent evaluations placed it #4 of 189 on the AA Intelligence Index.
  • DeepSeek introduced the first time-of-day surge pricing from a major model API (2x listed rates during Beijing peak hours, effective Jul 24 alongside its legacy-alias retirement) — importing electricity-market demand management into inference pricing.
  • The counter-move came from a US lab: Thinking Machines shipped Inkling under clean Apache 2.0 with weights live on day one — while K3's weights remained promissory (license unpublished, HF repo 404, AA classifying it proprietary until the Jul 27 drop).

Sources Kimi K3 pricing docs and launch analysis · DeepSeek surge pricing and alias retirement notice · AA Intelligence Index K3 placement · Inkling Apache 2.0 day-one release

Abductive · 66% confidence

Compute is becoming a traded commodity across former battle lines — the market structure is shifting from vertical exclusivity to a lease-and-credit market for GPU capacity, with credit and equity telling opposite stories at the most contract-concentrated operator. Resolution point: CoreWeave's early-August Q2 print.

Steel-man: The divergence is not sector-wide: Nebius equity ROSE ~8% on its debt announcement — the market read that facility as validating, not contradicting, the equity story — and the CoreWeave drawdown figure is secondary-press, not filed. The claim survives at the single-name level (CoreWeave's rated debt vs its de-rated equity is a real, dated tension), but a sector-wide credit-vs-equity split is not yet in evidence; the Q2 prints adjudicate.

  • Anthropic entered reported talks to lease up to ~$10B of capacity from Meta — a top-2 lab renting a rival's fleet — days after Meta doubled Hyperion to >$50B, meaning surplus hyperscaler capex is being productized as a fifth cloud.
  • The debt market institutionalized in the same week: Nebius closed an SEC-filed $775M GPU/contract-backed facility at SOFR+2.50%, Fitch rated CoreWeave's $2.6B take-or-pay delayed-draw term loan at BB+, and Upper90 wrote the first inference-ASIC-collateralized loan.
  • At CoreWeave the two markets disagree: lenders underwrote its contracted cash flows at investment-adjacent spreads the same week its equity sat ~35% below the Meta Compute report level — both cannot be right at current spreads.

Sources Anthropic-Meta compute lease talks · Nebius $775M secured facility (6-K) · Fitch CoreWeave DDTL rating note

Deductive · 79% confidence

The binding constraints on AI deployment have fully migrated downstream of silicon and capital — confirming the framework's Hypothesis 5 at a new layer: in one week, the world's foundry said packaging caps its customers' growth, the largest US grid cleared short at an administratively capped price, and a state added the first policy-layer pause — while capital itself flowed unimpeded. Resolution points: the September PJM backstop auction and TSMC's Q3 packaging-capacity commentary.

Steel-man: Each leg has a softer reading: TSMC's capex raise can be read as still-conservative risk transfer (Stratechery's argument — the foundry deliberately underbuilds and lets customers eat shortage risk as foregone revenue, so 'packaging-constrained' partly reflects choice, not just physics); the PJM cap is a policy instrument that FERC can lift; and EO 62 landed a month after the legislature passed a stricter 20 MW bill the governor is expected to veto — arguably a softening maneuver, not an opposition victory. The claim survives because all three softer readings still leave the constraints downstream of silicon and capital — they dispute the constraints' permanence, not their location.

  • TSMC posted a record quarter, raised capex to $60-64B, and added $100B to Arizona — capital is abundant — while CEO C.C. Wei stated advanced packaging capacity is 'so tight that now it's limiting my customers' growth.'
  • PJM's 2028/29 auction cleared at the FERC cap for the third straight year, 6,831 MW short of the reliability requirement, with PJM's own uncapped simulation 71% higher — administrative price suppression masking a worsening shortage (dollarized in this issue's house measurement at ~$11.6B/yr).
  • New York signed EO 62 the next day — the first statewide data-center permit moratorium — adding a political layer on top of the physical and administrative ones, per Hypothesis 5 (power as the binding constraint).

Sources TSMC Q2 6-K and earnings call (packaging commentary) · PJM 2028/29 BRA results report · New York EO 62 text

Synthesis · Thesis test

Hypothesis 1 · Supported

The cycle is accelerating, not slowing.

Two trillion-parameter-class open MoEs landed in a single week (K3 at a vendor-reported 2.8T, Inkling at 975B), and an open-weights-slated model reached within ~3 points of the closed frontier on the independent index just five weeks after GPT-5.6's government preview began — a gap-closing cadence that took DeepSeek V4 months at the start of the year. On the hardware clock, TSMC's 2nm hit its first material revenue quarter (3% of wafer revenue) while FY26 growth guidance was raised to 'slightly above 40%', and Intel shipped the first high-volume High-NA EUV logic. Both the capability clock and the silicon clock ticked faster this week.

Counter-evidence: Gemini 3.5 Pro slipped a third time past its leaked Jul 17 target — a top-3 lab repeatedly missing its own frontier cadence is the week's clearest deceleration signal, and TSMC's packaging warning caps how fast the silicon clock can convert to shipped racks.

Sources K3 independent frontier-tier placement · TSMC Q2: 2nm at 3%, guidance raised · Intel High-NA EUV production milestone

Hypothesis 2 · Supported

Capital is concentrated, returns are diffuse.

The realized profits again landed at the physical layer while the model layer spent: TSMC booked a record $40.2B quarter with net income up 77% YoY (EPS +74.5%) and ASML beat the high end of guidance with 2027 EUV nearly fully booked — audited, banked returns — while frontier labs cut effective prices, rented rivals' compute, and Meta committed >$50B to a single site with no revenue attached. The concentration side sharpened too: two single-company sites (Hyperion, Tembo) now each approach the aggregate size of Japan's entire national AI program, and the clearest new revenue stream in the category was lease payments and debt coupons on contracted GPU capacity, not model margins.

Counter-evidence: SemiAnalysis's pre-IPO model puts Anthropic at >$1B quarterly operating profit (analyst estimate, grade 2) — if the filing confirms it, returns are beginning to concentrate at the model layer too, which would blunt the 'diffuse returns' half of the hypothesis.

Sources TSMC record Q2 (SEC-filed) · Meta Hyperion >$50B expansion

Hypothesis 3 · Supported

Networking is the durable layer.

The supply side voted with capital this week: Tower Semiconductor committed ~$3B (state-co-funded) to silicon photonics capacity explicitly for optical interconnect demand, Microsoft became the first hyperscaler to deploy 3M's expanded-beam optical connectivity across AI data centers (attacking deployment velocity, not just link speed), and Japan's sovereign program standardized on a single-vendor Ethernet fabric — a state-scale Metcalfe machine. The revenue-side test (interconnect growth outpacing compute growth at colo operators) comes with the Q2 prints starting Jul 29 and is now prediction p64.

Counter-evidence: This week's support is all supply-side and announcement-grade: no new fabric-protocol spec shipped (UEC has sat at v1.0.2 since January), and the hypothesis's actual falsification test — interconnect revenue growth vs compute revenue growth — has no fresh data until the Jul 29+ prints. Capital commitments to optics capacity are conviction, not confirmation.

Sources Tower $3B silicon photonics expansion · 3M-Microsoft EBO deployment partnership

Hypothesis 4 · Strained

Open weights pull the floor up.

An honest week for this hypothesis is a strained one. The supporting half is real: Inkling shipped Apache 2.0 weights at 975B scale with day-0 serving support and a calibrated NVFP4 checkpoint that fits ~600GB, explicitly positioned as a fine-tuning base for private deployment. But the week's flagship 'open' release cut against the mechanism on both of the hypothesis's load-bearing words: K3 shipped NO weights (hosted-only, license unpublished, HF repo 404) and raised rather than lowered the price floor (~3x its predecessor, closed-vendor rate card). A hypothesis about open weights pulling the floor UP cannot count a weights-less launch that pulled prices up as support. If the Jul 27 drop lands with a clean commercial license, the verdict likely reverts to supported — with the sharpened form 'open weights re-route demand to self-hosting even as open API prices converge upward.' If the weights slip or the license restricts, this hypothesis needs revision in thesis.ts.

Counter-evidence: The strain is the verdict this week; the counter-counter case is Inkling itself (the mechanism operated cleanly at 975B scale from a US lab) plus Moonshot's 64+ accelerator supernode self-hosting guidance, which presumes the weights really are coming.

Sources K3 hosted-only launch, license unpublished · Inkling Apache 2.0 release with NVFP4 checkpoint

Hypothesis 5 · Supported

Power is the binding constraint for the next 24 months.

The strongest single-week confirmation since the hypothesis was authored — which, after five consecutive supported weeks, is itself worth noting: this hypothesis is approaching consensus and may soon stop differentiating (a sharpening candidate for thesis.ts). The evidence: PJM cleared at the administrative cap for the third consecutive year while coming up 6,831 MW short — with the operator's own uncapped simulation 71% above the clearing price and only 525 MW of new generation clearing — and New York layered the first statewide permit moratorium on top of queue physics the very next day. The constraint now has three stacked layers: physical (generation shortfall), administrative (price caps masking scarcity), and political (state-level pauses).

Counter-evidence: The political layer is contested, not one-way: EO 62 is narrower than headlines suggest and arguably a softening maneuver against the legislature's stricter 20 MW bill (veto expected). And the constraint's binding-ness is partly administrative choice — FERC lifting the collar or the September backstop auction clearing new generation would relieve price pressure without any new physics.

Sources PJM 2028/29 BRA results (at cap, 6,831 MW short) · New York EO 62 statewide moratorium

Synthesis · Pattern watch

Inductive · 3 weeks observed

Model-API pricing is converging from both directions onto a $1-3 input / $6-15 output frontier band — closed labs cutting down into it, Chinese open-weight labs raising up into it.

Next expectation: DeepSeek V4's official GA pricing (due within days) lands at or above current V4-Pro rates rather than resetting the old ultra-cheap floor. If V4 GA undercuts to pre-K3 Chinese pricing levels, the convergence read fails and the price war resumes.

  • W27: cost-per-task became the explicit competitive axis as closed-lab pricing moves anticipated GPT-5.6 tiering.
  • W28: the closed frontier repriced downward — Terra GA at $2.50/$15 (half of GPT-5.5) and Grok 4.5 at $2/$6.
  • W29: the open-weight side repriced upward — Kimi K3 at $3/$15 (roughly 3x its predecessor, parity with Sonnet 5) and DeepSeek adding 2x peak-hour surcharges from Jul 24.

Inductive · 5 weeks observed

Government action is a standing gate on frontier-model availability — and the gate now includes labs gating themselves.

Next expectation: The GPT-Red technical preprint ships within two weeks but the model does not — and no API access materializes. If OpenAI releases GPT-Red weights or endpoints in any form, the self-gating leg of this pattern breaks.

  • W25: Fable 5 and Mythos 5 suspended under US export controls (Jun 12).
  • W26: Mythos 5 restored only for ~100 'Annex A' critical-infrastructure organizations.
  • W27: GPT-5.6 previewed to ~20 government-vetted partners at the US government's request.
  • W28: GPT-5.6 GA'd only after a 12-day government-coordinated preview with CAISI evaluations; Beijing rationed H200 purchases and pushed the Manus unwind.
  • W29: OpenAI announced GPT-Red will never be released due to offensive capability — the first explicit capability self-gating of a disclosed frontier-scale model — while Gemini 3.5 Pro slipped a third time.

Inductive · 5 weeks observed

Physical-input scarcity (memory, power, now packaging) keeps marking itself to market with progressively harder money and harder words.

Next expectation: SK hynix (Jul 29) and Samsung (Jul 30) Q2 prints disclose HBM substantially sold out or committed into 2027 with capex raises attached. Prints showing softening HBM commitments or flat memory capex would be the first hard counter-evidence in five weeks.

  • W25-W26: all three HBM makers volume-shipping HBM4; Micron confirmed 2026 supply fully contracted.
  • W27: SK hynix filed a ~$29.4B Nasdaq listing; price caps reportedly removed from long-term memory contracts.
  • W28: the IPO closed at $26.5B with 7x demand; Samsung guided to a record quarter; Meta and MARA paid premiums for secured power.
  • W29: PJM cleared at the cap and 6,831 MW short; TSMC's CEO said packaging is 'limiting my customers' growth' while raising capex to $60-64B; server DRAM contracts guided +13-18% QoQ for Q3; Tower committed $3B to optics capacity.

Synthesis · Second-order effects

Q4 2026 - H1 2027

Meta doubles Hyperion to >$50B while simultaneously negotiating to lease capacity to Anthropic, as GPU-backed credit institutionalizes at rated spreads.

A fifth cloud priced against neoclouds emerges from surplus hyperscaler capex, splitting the neocloud category by funding cost: operators with bankable take-or-pay contracts refinance into cheap secured debt while the rest face equity markets that just repriced their concentration risk. Expect consolidation — the weaker half of the neocloud category becomes acquisition inventory for hyperscalers and infrastructure funds within two or three quarters.

Who moves
Neocloud equity holders and lenders, frontier labs negotiating compute, infrastructure funds, and hyperscaler corp-dev teams

2027

New York signs the first statewide data-center permit moratorium one day after PJM clears at the cap with a 6,831 MW shortfall.

Data-center opposition movements in states with real AI load now have a governor-signed template, converting diffuse local resistance into replicable state policy — which raises the siting-risk premium on any campus without secured permits, accelerates the behind-the-meter generation and scale-across fabric workarounds, and pulls forward the value of already-permitted powered land as a distinct asset class.

Who moves
Data-center developers without generation balance sheets, state energy regulators, utilities, and holders of permitted powered-land inventory

30-90 days

Kimi K3 reaches the closed-frontier tier on independent evaluations with open weights promised July 27 and self-hosting guidance targeting 64+ accelerator supernodes.

If the weights and a commercial-friendly license land, data-residency-constrained enterprises and sovereign programs get their first credible exit from frontier API dependency — which re-routes inference spend from closed-lab APIs toward private accelerator clusters, HBM, and interconnect, and forces closed labs to defend on harness, tooling, and injection-hardening (exactly where OpenAI positioned GPT-Red this week) rather than raw capability.

Who moves
Closed-lab API revenue, sovereign AI programs, enterprise infrastructure teams, and the accelerator/HBM supply chain

Synthesis · Strategic outlook

This week moved the 12-month posture on three fronts. First, procurement: the open-vs-closed decision is no longer about a capability gap (~3 Index points) or a price gap (K3 on the closed vendors' rate card) — it is about deployment control, and July 27 is the date that tells you whether frontier-class self-hosting is real; budget the evaluation now, because if the weights land clean, every closed-lab contract renewal after August negotiates against a credible internal alternative. Second, capital structure: compute is becoming a two-sided market — leased across rival lines, financed with rated secured debt — so treat GPU capacity commitments the way treasurers treat interest-rate exposure: as a portfolio of owned, leased, and optioned positions rather than a single vendor bet. Third, the constraint stack: with packaging capped at the foundry, power capped at the auction, and permits now pausable by executive order, the scarce inputs for 2027 are all downstream of money — which means the durable advantages are secured packaging allocation, energized land, and permitted sites, and the capital markets just started pricing all three accordingly.

Where we differ

Extend

Kimi K3 marks 'the end of super cheap Chinese AI' — Chinese labs have stopped competing primarily on price (a launch-day frame Simon Willison echoed: the most expensive Chinese-lab model to date).

Right, and it is half the picture: the closed labs cut DOWN into the same band last week (Terra at half GPT-5.5, Grok 4.5 at $2/$6). The two-sided convergence onto one $1-3/$6-15 flagship band is the market-structure event, it shifts differentiation to who actually ships weights — and it is falsifiable: if DeepSeek V4's GA pricing resets the ultra-cheap floor, the convergence read fails (prediction p66, 84%).

Sources the-decoder, Jul 16

Extend

PJM's capped clearing price hides widening scarcity: the capped-vs-simulated gap has grown 18% -> 59% -> 71% across three auctions — administrative suppression is widening, not easing.

We borrow the gap trendline with credit and add the number nobody printed: dollarized, the suppression is ~$11.6B/yr RTO-wide, ~$84,000 per MW-year, an $8.4M/yr hidden premium on an illustrative 100 MW AI load — see this issue's house measurement for the arithmetic and its caveats.

Sources Modo Energy, ~Jul 14

Open

TSMC deliberately underbuilds and offloads shortage risk onto customers as foregone revenue — so a capex raise is risk transfer, not pure demand confirmation. (A standing January argument we are applying to this week's raise, not a this-week take.)

Partially persuaded: Wei's packaging commentary reads as physics, but Thompson's lens explains why the constraint persists through record capex — underbuilding is a choice that transfers shortage risk downstream. We carry it as the steel-man on our constraint-stack claim and will treat the new fabs' groundbreaking dates and tool orders as the adjudicating evidence.

Sources Stratechery, Jan 26

Differ

New York's EO 62 is a first-in-the-nation moratorium and a potential turning point against data-center development.

The scope is narrower than the headlines (discretionary state environmental permits, ≥50 MW, completed applications exempt) and the timing cuts the other way: the order landed a month after the legislature passed a stricter 20 MW bill the governor is expected to veto — arguably a softening maneuver against the legislature, not an opposition victory. Precedent matters; crackdown framing doesn't hold.

Sources CNBC / law-firm client alerts, Jul 14-15

Open

'Just like DeepSeek,' Kimi K3 is forcing Western labs to question their compute advantage — China may have caught the frontier.

The week's most strategically loaded open question, and we decline to call it: a vendor-reported 2.8T-parameter model trained under export controls either falsifies the compute-poverty thesis or the parameter count. The Jul 27 technical report and independent replications of the vendor-harness benchmarks are the evidence that decides; until then both triumphalist and dismissive reads are ahead of the data.

Sources the-decoder, ~Jul 17

Levers

MetricCurrentPriorDirectionThreshold
Frontier lab cash position (avg months runway, disclosed-burn labs)~30-40 mo (range; unaudited inputs); no new primary capital — the movement was compute-sourcing (Anthropic-Meta lease talks, reported ~$10B/2yr, disputed)~34-37 mo; Anthropic at implied $1.2T on secondaries (broker-reported), IPO calendars unchangedflat<18 mo triggers re-rating risk
Hyperscaler capex / AI revenue ratio (top 4 weighted)~5.0-5.5; Meta Hyperion doubled to >$50B/5 GW, Google behind 2.7 GW Wyoming 'Project Tembo' — numerator still hardening ahead of the Jul 22-30 earnings wave~5.0-5.3; Meta targets 14 GW of compute in 2027 (2x 2026) on $125-145B capex guidanceup>6.0 invites investor pushback at next earnings
CoreWeave revenue backlog$99.4B as of Mar 31 (+284% YoY), restated in Fitch's Jul 16 DDTL note; equity down ~35% since the Meta Compute report~$100B reported; Helios Phase I (133 MW) delivered on schedule — backlog now converting to lease revenueflatConversion velocity matters more than gross figure
NVIDIA Q-over-Q data center revenue$75.2B Q1 FY27 (unchanged); Huang in Tokyo: 'Vera Rubin is already in production' — delay narrative rebutted, but no customer-delivery date named$75.2B Q1 FY27; Q2 guide $91B (reports Aug 26); SemiAnalysis sees H2 ~20% above consensus despite Kyber disputeflatQ2 FY27 guide $91B implies further +21% QoQ
Open vs closed gap on coding (SWE-Bench / agentic)~3 pts on the AA Intelligence Index — Kimi K3 at 57.1 vs Fable 5's 59.9 — with K3 weights pending Jul 27; K3 already #1 on LMArena Frontend Code ArenaEffectively closed on cost-quality: GLM 5.2 statistically tied with Opus 4.8 at $1.28 vs $1.94/task (Databricks); Hy3 adds Apache-2.0 agentic-search leaddownSustained open lead reshapes enterprise procurement
Sovereign AI commitments (count / aggregate $)~15 / ~$186B; Japan FRONTia adds ¥1T (~$6.2B/5yr) — the first national program explicitly scoped to physical AI and robotics~14 / ~$180B+ (flat; Meta Alberta and MARA Texas are corporate capital on power-rich land, not sovereign programs)up
PJM 2026/27 capacity auction price ($/MW-day)$325.00 — the 2028/29 BRA cleared at the (lowered) FERC cap Jul 14, 6,831 MW short of the reliability requirement; PJM's uncapped simulation: $554.72 RTO / $776.69 ComEd$329.17; 2028/29 BRA bids closed Jul 7 — results post Jul 14 after 4 p.m. ET (prediction p54 resolves)flat11x in 24 months — power is the new binding constraint
Time-to-power, busiest US markets (months)60-84 unchanged; New York adds the first statewide ≥50 MW discretionary-permit moratorium — state policy now stacks on top of queue physics60-84; hyperscalers routing around queues — Meta fully funds its own generation in Alberta because the grid cannot host multiple large loadsflat
Cost-per-task, frontier reasoning modelFloor $0.04/task (DeepSeek V4 Pro); median across this week's named frontier models ~$0.63 (V4 Pro $0.04, GLM-5.2 $0.32, K3 $0.94, Sol $1.04; basis: AA Intelligence Index per-task) — and DeepSeek adds the first peak-hour surge pricing~$0.06-$0.12 effective; GPT-5.6 GA tiering (Sol $5/$30 / Terra $2.50/$15 / Luna $1/$6), Grok 4.5 at $2/$6 — and harness choice swings per-task cost 2x+flat
Custom silicon share of incremental AI compute~34-37% unchanged; Google reportedly pitching TPUs to NVIDIA-centric neoclouds with financial backstops — custom silicon escalating from internal cost play to merchant distribution fight~34-37%; Meta's Iris enters production in September, Broadcom-Apple extended through 2031 (8-K), AWS raising Trainium 3 orders 20-30%flat>35% materially compresses merchant GPU pricing

Frontier lab cash position (avg months runway, disclosed-burn labs)

Methodology v2 (see METHODOLOGY.md): basis narrowed to disclosed-burn labs (OpenAI, Anthropic, xAI) — Google DeepMind is excluded because a subsidiary funded from Alphabet's balance sheet has no meaningful standalone runway, and averaging it in produced a number that meant nothing. Displayed as a range because both inputs (cash position, net burn) are unaudited. No primary financing closed in-window; SoftBank's second $10B OpenAI tranche executed Jul 1 (pre-window). The structural shift is labs negotiating compute from rivals' fleets — Anthropic-Meta talks on top of the earlier $45B Colossus 1 commitment — which converts capex exposure into opex flexibility. Boards should read this as runway-preserving behavior, not distress.

Hyperscaler capex / AI revenue ratio (top 4 weighted)

Top 4 hyperscalers (MSFT, GOOG, META, AMZN) weighted aggregate of capex divided by AI-attributable revenue. Two single-site adds this week (Hyperion >$50B, Tembo ~$50B reported) push the committed numerator up without any new revenue disclosure — the ratio is unmeasurable to audit standard because AI revenue isn't segmented. The Jul 22-30 earnings wave (Alphabet, Microsoft, then Meta) is the first test of whether guides absorb these adds or raise.

CoreWeave revenue backlog

Booked but unrecognized revenue; official next print is Q2 in early August. This week the backlog got third-party validation of a different kind: Fitch rated a $2.6B contract-backed DDTL at BB+ against take-or-pay contracts, meaning a rating agency underwrote the backlog's bankability even as public equity de-rated it. Watch the Q2 print for anchor-customer mix and conversion velocity — credit and equity cannot both be right at current spreads.

NVIDIA Q-over-Q data center revenue

No earnings event in-window. Huang's Tokyo statement rebuts the early-July delay reporting for NVL72 racks, and Japan's FRONTia order (27,500 Rubin GPUs) puts a named sovereign customer behind the ramp — but TSMC's same-week warning that packaging capacity is 'limiting my customers' growth' is the supply-side caveat that keeps the delivery-date question open until the Aug 26 print.

Open vs closed gap on coding (SWE-Bench / agentic)

The absolute-capability gap narrowed to ~3 Index points this week — the closest an open-weights-slated model has come to the closed frontier — and K3 topped a flagship human-preference board outright. Two caveats keep this from resolving: K3 is hosted-only until Jul 27 (AA classifies it proprietary in the interim), and its vendor-reported coding numbers ran on its own KimiCode harness. If the weights and license land clean, this lever's threshold — sustained open parity reshaping procurement — is effectively triggered.

Sovereign AI commitments (count / aggregate $)

Analyst-curated count of sovereign/national AI-compute commitments. FRONTia is the template evolution to watch: consortium structure (48 firms), national data, a hard chip order (27,500 Rubin GPUs at 140 MW), and a single-vendor fabric — sovereign programs are becoming reference-design purchases rather than bespoke buildouts. NVIDIA's SEC-filed commentary already attributes half its data-center revenue to non-hyperscale demand including sovereign; expect FRONTia clones in other industrial states within two quarters.

PJM 2026/27 capacity auction price ($/MW-day)

Third consecutive auction pinned at the administrative cap; the nominal 2.5% decline is entirely the cap formula, not loosening scarcity. The information is in the gap: the capped-vs-simulated spread widened from 18% to 59% to 71% across three auctions (trendline credit: Modo Energy), only 525 MW of new generation cleared, and total cost still hit $16.4B. Budget at-cap through 2028; the September backstop auction and PJM's data-center connect-and-manage FERC filings are where the real price of AI load gets set.

Time-to-power, busiest US markets (months)

Months from new-load interconnection request to energization. No fresh queue data in-window, but the constraint acquired a policy layer: EO 62 pauses discretionary environmental permits for ≥50 MW facilities statewide while New York writes a generic environmental impact statement. Narrow in operational scope, precedential in kind — developers should now underwrite state-policy pause risk in any market with active data-center opposition, on top of the 60-84 month grid reality.

Cost-per-task, frontier reasoning model

DeepSeek's 2x Beijing-peak-hours surcharge (from Jul 24) imports electricity-market demand management into model APIs

Methodology v2 (see METHODOLOGY.md): this lever previously conflated the cheapest model's cost (a floor) with a median label; it now reports both explicitly on a stated basis (AA Intelligence Index per-task this week). The floor held at $0.04 while two structural changes landed: K3's always-on max-effort thinking shows sticker price diverging from effective cost in the expensive direction (130M tokens on the AA eval vs 63M peer average), and DeepSeek's time-of-day surge pricing means per-task cost now varies by clock. Cost models should carry cache-hit rates, reasoning-effort settings, and now time-of-day as explicit variables.

Custom silicon share of incremental AI compute

Estimated share of incremental AI compute capacity on hyperscaler custom silicon. No confirmed share shift this week — Nebius, Lambda, and CoreWeave all publicly demurred on the reported Google TPU pitch — but the pitch itself (backstopping TPU data centers and renting capacity back) is the clearest signal yet that custom silicon is contesting NVIDIA's own neocloud channel rather than just displacing internal purchases. NVIDIA reportedly counter-offering incentives to Nscale confirms the channel fight is live.

Predictions

Software · 72% confidence

Moonshot publishes Kimi K3 open weights on Hugging Face with a license permitting commercial self-hosting by August 10, 2026 (vendor-committed 'by July 27', with slippage buffer).

ID
p61-kimi-k3-weights-aug10
Deadline
By August 10, 2026
Trigger
Kimi K3 weights live on Hugging Face with published license text; Artificial Analysis reclassifies K3 from proprietary to open-weights.

Software · 76% confidence

DeepSeek retires its legacy deepseek-chat and deepseek-reasoner aliases on July 24 as scheduled and ships DeepSeek V4 to official GA by July 31, 2026, with peak-hour surge pricing in effect.

ID
p62-deepseek-v4-ga-jul31
Deadline
By July 31, 2026
Trigger
DeepSeek API docs showing V4 GA model IDs and the alias-retirement notice executed; surge pricing live in the rate card.

Hardware · 68% confidence

SK hynix's July 29 Q2 earnings call discloses that 2027 HBM capacity is substantially sold out or committed under long-term agreements, extending the memory-scarcity trade into a second year.

ID
p63-hbm-soldout-2027
Deadline
By July 29, 2026
Trigger
SK hynix Q2 2026 earnings call commentary on 2027 HBM capacity commitments and capex.

Networking · 71% confidence

The largest colocation operators' Q2 prints (starting July 29) show interconnect/fabric revenue growth again outpacing overall revenue growth, sustaining the Metcalfe read for a second consecutive quarter.

ID
p64-colo-interconnect-outpaces
Deadline
By August 15, 2026
Trigger
Q2 2026 colo earnings disclosures: interconnection and fabric revenue growth rates vs total revenue growth.

Power · 57% confidence

At least one additional US state announces a moratorium, discretionary-permit pause, or equivalent statewide restriction on large data-center development by October 31, 2026, following New York's EO 62.

ID
p65-state-moratorium-copycat
Deadline
By October 31, 2026
Trigger
A governor's executive order or enacted state legislation pausing or restricting ≥50 MW-class data-center permitting in a second state.

Software · 84% confidence

DeepSeek V4's official GA pricing does not reset the ultra-cheap floor: off-peak deepseek-v4-pro output pricing stays at or above ¥6 (~$0.85) per MTok through August 31, 2026 — the kill-condition test for this issue's price-band-convergence claim.

ID
p66-no-cheap-floor-reset
Deadline
By August 31, 2026
Trigger
DeepSeek's published API pricing page (api-docs.deepseek.com) for the GA deepseek-v4-pro model, off-peak output rate.

Prior predictions scored

Pending · Software

Gemini 3.5 Pro reaches public general availability — a callable API model ID with published pricing — by July 31, 2026, after slipping past its June window and the reported July 17 target.

Missed the leaked Jul 17 target this week — its third slip — with reporting citing hallucination/reliability gaps and possible stopgap releases; the public API still lists only gemini-3.5-flash and gemini-3.1-pro-preview. Two weeks remain on the deadline; prediction markets price ~81% by Jul 31.

ID
p57-gemini-3-5-pro-ga-jul31
Confidence
58%
Deadline
By July 31, 2026
Trigger
Google Gemini API model list / pricing page showing a GA gemini-3.5-pro model ID.

Pending · Software

At least one major agent platform (OpenAI, Anthropic, GitHub, or Cursor) ships product-level per-task or per-harness cost telemetry or routing controls — beyond session budget caps — by August 31, 2026.

Adjacent evidence accumulated: Atlassian shipped DX AI cost-vs-output measurement (Jul 15) and Claude Code v2.1.212 shipped session-level search/subagent budgets — but Atlassian is outside the four named platforms and session caps are explicitly excluded by the prediction's own terms. The pattern is advancing; the specific trigger hasn't fired.

ID
p58-harness-cost-telemetry
Confidence
64%
Deadline
By August 31, 2026
Trigger
Product changelog or GA announcement exposing per-task cost measurement or harness-level cost controls.

Partial · Hardware

TSMC's July 16 Q2 earnings raise or reiterate the top end of full-year 2026 capex guidance and report HPC/AI platform revenue up more than 50% year over year, confirming the packaging-constrained AI capex ramp.

The capex leg hit decisively — TSMC raised the entire FY26 range to $60-64B (from $52-56B), well beyond reiterating the top end. The revenue leg fell just short: HPC at 66% of a $40.2B quarter implies ~47% YoY platform growth versus the >50% bar. The thesis behind the prediction (packaging-constrained ramp) was confirmed on the record by the CEO's packaging warning (quoted in the hardware lens).

ID
p59-tsmc-q2-capex-raise
Confidence
62%
Deadline
By July 16, 2026
Trigger
TSMC Q2 2026 earnings release and investor call (Jul 16; June revenue print Jul 13, typhoon-delayed).

Pending · Networking

A second named vendor or operator announces a commercial cross-data-center scale-across AI fabric deployment or product launch — following DriveNets/WhiteFiber — by September 30, 2026.

No second vendor announcement in-window; DriveNets/WhiteFiber drew continued trade analysis and their Q3 commercial launch remains on track. Hot Interconnects 2026 (Aug 19-21, themed 'Scale-Up, Scale-Out, Scale-Across') is the likely venue for the next entrant.

ID
p60-scale-across-follow-on
Confidence
61%
Deadline
By September 30, 2026
Trigger
Vendor or operator press release for a commercial (not lab) multi-site training-fabric deployment; WhiteFiber's own Q3 commercial launch also qualifies if it lands with a named second customer.

Watchlist

Jul 20-24

OpenAI's GPT-Red technical preprint

Promised 'later this week' — enables third-party scrutiny of the 84%-vs-13% red-teaming and 0.05% injection-failure claims that are currently vendor-reported only. If it substantiates, injection-hardening becomes a hard procurement criterion industry-wide.

Jul 22-23

Q2 earnings triple-header: ServiceNow and Alphabet (Jul 22), then SAP, Intel, and AMD's Advancing AI event (Jul 23)

First reads on agentic-AI monetization (ServiceNow), cloud AI revenue (Alphabet), the closed-gateway Joule strategy (SAP), 18A economics post-High-NA (Intel), and MI455X/Helios rack shipment clarity (AMD).

Jul 24

Two hard deadlines: M365 Copilot's OpenAI-subprocessor setting auto-enables, and DeepSeek retires its legacy API aliases

M365 tenants that don't opt out are auto-enrolled in a changed data-processing chain (silence is consent); DeepSeek's alias retirement plus surge pricing is a forced-migration event for one of the most-called model APIs.

Jul 27

Kimi K3 open-weights drop, license text, and technical report

The single highest-information event in view: if the weights and a commercial-friendly license land, frontier-class capability becomes self-hostable for the first time and prediction p61 resolves early; if they slip or the license restricts, the 'open frontier' narrative resets.

Jul 29-30

Memory and hyperscaler prints: SK hynix Q2 (Jul 29), Microsoft FY26 Q4 (Jul 29), Samsung divisional results (Jul 30), colo interconnect Q2 prints begin (Jul 29)

The supply-side read on 2027 HBM commitments (p63), whether hyperscaler capex guides absorb this week's Hyperion/Tembo adds, and the interconnect-revenue test of the networking hypothesis (p64).

Changelog

  • Kimi K3 and Inkling added to the LLM Evolutionary Tree this week — see the Model Pulse treeDelta for placement detail.
  • W29-r2 revision (Jul 18, post-review): a four-persona editorial review of the first W29 publish drove structural changes now standard for every issue — steel-man fields on every synthesis connection, named counter-evidence on every thesis verdict, a house measurement section (one number we compute ourselves), a 'Where we differ' box against the top tier's takes, and the full prediction track record with Brier scoring. Content corrections in this revision: the 'Sonnet parity' claim now notes Sonnet 5's live $2/$10 intro price (K3 is ~50% above street); the EO 62 read adds the preempted stricter 20 MW legislative bill; the credit-vs-equity connection is narrowed to CoreWeave (Nebius equity rose on its debt deal); Hypothesis 4 is honestly scored 'strained' pending the Jul 27 K3 weights drop.
  • Metrics methodology v2 (see content/industry/METHODOLOGY.md): the frontier-lab runway lever now covers disclosed-burn labs only (Google DeepMind excluded — a subsidiary has no standalone runway); the cost-per-task lever reports floor and median separately on a stated basis; burn-to-revenue is now uniformly capitalIn/revenueOut (the Frontier Labs row corrects from ~1.3x to ~4.5x under the consistent definition — prior issues are not restated).