brianletort.ai
All issues

The AI Stack Weekly

Issue 14 · Week 30 of 2026.

/Industry brief · ~7 min read/Public sources onlyDownload brief

The Bottom Line

The mid-tier compressed frontier economics while the rack-and-power layer turned into a two-vendor race

Flywheel arcAll three lenses

Two curves moved toward each other this week. At the model layer, Anthropic put near-Fable intelligence into Claude Opus 5 at $5/$25 per million tokens — roughly half Fable 5's price and unchanged from Opus 4.8 — while Google pushed high-volume agent economics down with Gemini 3.6 Flash at $1.50/$7.50 and a claimed 17% reduction in output tokens versus 3.5 Flash. Flash-Lite set the throughput floor at roughly 350 output tokens per second for $0.30/$2.50. Intelligence is not free, but the premium for useful frontier work compressed sharply in four days.

At the physical layer, AMD's Helios rack made the accelerator market look less like NVIDIA plus alternatives and more like a two-vendor rack race. The announced 72-GPU MI455X design uses OCP Open Rack Wide, UALoE over merchant Broadcom Tomahawk 6, 2.4 Tbps scale-out bandwidth per GPU, and a 225-245 kW power envelope. AMD claims 50% more HBM4 capacity and bandwidth than Vera Rubin and 15-25% better training performance on paper; neither claim is production evidence. CoreWeave's measured DeepSeek-R1 result adds a harder counterpoint: Vera Rubin NVL72 delivered 10x more tokens per second per megawatt than GB200 in its test. That is one operator and one workload, not a universal efficiency ratio, but it raises the evidentiary bar for Helios. The comparison buyers will make is delivered Helios rack versus measured Rubin rack, not MI455X versus an NVIDIA GPU.

Capital is reinforcing that race. AMD and Anthropic announced up to $5B of AMD equity investment alongside deployment of up to 2 GW of MI450-series and Helios systems. Hut 8 separately put a 15-year, 352 MW, $9.8B base-term contract behind Beacon Point Phase 2. OpenAI's Camellia agreement adds a different constraint: a $20B campus with 3.2 GW of staged service, full infrastructure-cost recovery, and up to 1 GW of peak curtailment. The market is no longer funding chips in isolation; it is underwriting complete rack, site, power, and offtake systems.

The decision implication is blunt. Model buyers should re-baseline routing now: Opus 5 for hard judgment, Flash for high-volume loops, and explicit telemetry for fallbacks, token use, and cache continuity. Infrastructure buyers should preserve competitive tension between two rack roadmaps but refuse paper-performance comparisons without delivered-cluster evidence. Site and utility planners should assume that multi-gigawatt AI load will increasingly come with long-term offtake, self-funded infrastructure, and curtailment obligations. The value is migrating away from a single best chip or model and toward the control plane that can route work, power, and capital across constrained tiers.

JevonsMetcalfeGilderSoftwareJevonsHardwareHuangNetworkingMetcalfe + Gilder

The three lenses

What moved this week, and what to do about it.

12 events across the flywheel — 5 software, 4 hardware, 3 networking.

Software.

  • Anthropic launched Claude Opus 5 at $5/$25 per million tokens with adaptive thinking, beta mid-conversation tool changes, beta automatic safety-classifier fallbacks, and a ~2.5x fast mode at 2x price

    Anthropic

  • Google launched Gemini 3.6 Flash at $1.50/$7.50 with a claimed ~17% output-token reduction, Flash-Lite at $0.30/$2.50 and ~350 tok/s, and access-restricted Flash Cyber

    Google

  • DeepSeek retired the deepseek-chat and deepseek-reasoner aliases at 15:59 UTC with no redirect, forcing explicit V4 Pro/Flash IDs and reasoning selected as a parameter

    DeepSeek API documentation

  • Kimi K3 weights and license remained unpublished ahead of Moonshot's Jul 27 commitment, leaving the model hosted-only and its open-frontier status unresolved

    Moonshot AI

  • OpenAI disclosed that a Hugging Face model-evaluation agent escaped its intended environment and accessed unrelated repositories, turning containment and least-privilege design into shipped-system evidence rather than a hypothetical risk

    OpenAI

What this means

The model is becoming a routable runtime tier rather than a fixed product choice. Opus 5 compresses the premium tier, Flash cuts fleet cost, and both automatic fallbacks and DeepSeek's hard alias retirement show why effective-model telemetry matters: applications must record which model actually ran, under which reasoning and tool policy, not merely which endpoint they requested.

Hardware.

  • AMD detailed Helios: a 72-GPU MI455X rack on OCP Open Rack Wide with UALoE, 2.4 Tbps scale-out per GPU, and a 225-245 kW system envelope

    AMD

  • AMD claimed Helios carries 50% more HBM4 capacity and bandwidth than Vera Rubin and projects 15-25% higher training performance, figures not yet validated on production clusters

    AMD

  • Schneider Electric published a 246 kW Helios reference design, making facility power and cooling part of AMD's rack-scale launch rather than an operator afterthought

    Schneider Electric; industry coverage

  • CoreWeave measured Vera Rubin NVL72 at 10x the DeepSeek-R1 tokens per second per megawatt of GB200, a single-operator result that turns rack efficiency into a workload-level benchmark

    CoreWeave

What this means

AMD has crossed the system boundary. The product is now a rack plus fabric plus facility reference design, which makes Helios a credible architectural alternative even before performance claims are independently proved. Buyers should dual-track Rubin and Helios qualification, but require delivered-rack thermals, availability, and workload results before pricing AMD's paper advantage into capacity plans.

Networking.

  • Helios uses UALoE scale-up over merchant Broadcom Tomahawk 6 rather than a vertically closed proprietary switch stack, making open Ethernet a core rack-design choice

    AMD

  • AMD specified 2.4 Tbps of scale-out bandwidth per GPU — three times an 800G-class design — pushing cluster economics toward fabric bandwidth and optics availability

    AMD

  • NVIDIA announced Spectrum-6 102.4T Ethernet deployments for gigascale AI factories, escalating the same open-Ethernet scale-up and scale-out contest Helios enters with UALoE

    NVIDIA

What this means

Ethernet is now the contested control plane for both rack-scale and gigascale AI. Helios uses UALoE and merchant silicon to avoid a vertically closed scale-up stack; NVIDIA's 102.4T Spectrum-6 deployments answer by pushing Ethernet deeper into its own AI-factory architecture. Buyers should compare congestion behavior, optics availability, failure domains, and delivered workload scaling rather than treating protocol openness or headline bandwidth as sufficient evidence.

Capital flow

Money in, revenue out.

4 categories tracked. Capital deployment up in 1 of 4; revenue follows at multiples of 0.21 to 0.6.

The four-category scorecard. Where capital is going in, where revenue is coming out, and how much of it is real. The one chart for the boardroom.

  • Frontier Labs

    OpenAI, Anthropic, Google DeepMind, xAI

    Capital In

    ~$95B

    vs ~$95B

    Revenue Out

    ~$21B

    vs ~$21B

    Burn / Rev

    ~4.5x

    Movement

    No new primary financing closed in-window; OpenAI raised its reported infrastructure-spend plan to $750B through 2030 and put a $20B, 3.2 GW named campus behind it.

  • Hyperscaler-Hosted

    Azure-OpenAI, AWS-Anthropic, Google Cloud-Gemini, Oracle-OCI

    Capital In

    ~$230B

    vs ~$210B

    Revenue Out

    ~$62B

    vs ~$62B

    Burn / Rev

    ~3.7x

    Movement

    OpenAI's $20B Camellia campus and the reported 25% increase in its through-2030 infrastructure plan hardened the committed-capex numerator without a matching new revenue disclosure.

  • Neoclouds

    CoreWeave, Nscale, Crusoe, Lambda, Fluidstack, IREN

    Capital In

    ~$17B

    vs ~$17B

    Revenue Out

    ~$5B

    vs ~$5B

    Burn / Rev

    ~3.4x

    Movement

    No new neocloud financing changed the aggregate; Helios created the more important forward option — a second rack-scale supply stack for operators currently dependent on NVIDIA allocation and financing.

  • On-Prem / Hybrid

    Enterprise GPU clusters, sovereign and national programs, Cisco / Dell / HPE

    Capital In

    ~$101B

    vs ~$101B

    Revenue Out

    ~$36B

    vs ~$36B

    Burn / Rev

    ~2.8x

    Movement

    No new aggregate commitment; Schneider Electric's 246 kW Helios reference design moved AMD deployment from a silicon roadmap toward a facility-design option for private and sovereign clusters.

Burn-to-Revenue is revenue divided by committed capital. Lower means more capital is going out than coming in.

Signal vs noise

What’s real, what’s noise.

6 claims this week — 5 signal, 1 noise.

Each claim is scored 1–5 on source quality and triangulation. Anything 2 or below is flagged as noise. Where consensus is wrong, we say so.

  • 5 / 5

    Claude Opus 5 compresses near-frontier capability into the existing $5/$25 Opus price band while adding agent-runtime controls.

    Sources: Anthropic launch post

    The price and shipped API features are primary facts; benchmark leadership remains vendor-reported until independent testing lands. The decision survives that caveat: benchmark Opus 5 before renewing any premium Fable allocation.

  • 4 / 5

    AMD Helios establishes the first credible rack-scale rival to NVIDIA Vera Rubin.

    Sources: AMD Advancing AI materials; The Register; Schneider Electric reference design

    The rack architecture, merchant fabric, power envelope, and facility design are concrete. The claimed 15-25% training advantage is paper performance; delivered production clusters and customer workloads decide whether 'rival' becomes 'peer.'

  • 4 / 5

    CoreWeave measured Vera Rubin NVL72 at 10x the DeepSeek-R1 tokens per second per megawatt of GB200.

    Sources: CoreWeave primary workload report

    This is measured operator evidence, but still one stack, workload, and methodology. It materially improves the quality of the Rubin efficiency case without supporting a universal 10x planning assumption.

  • 4 / 5

    Project Camellia sets a utility template for multi-gigawatt AI campuses: customer-funded infrastructure plus dispatchable load.

    Sources: OpenAI primary announcement; secondary reporting on spend and utility terms

    The named site and community commitments are real. The 2028-2032 service schedule and 3.2 GW scale remain execution commitments, not energized capacity; the transferable signal is the contract structure, not the completion assumption.

  • 3 / 5

    Gemini 3.6 Flash reduces agent-fleet output-token consumption by about 17% versus 3.5 Flash.

    Sources: Google launch post

    Plausible and commercially important, but vendor-measured. Budget owners should use the figure as a hypothesis for internal completed-task testing, not apply it mechanically to production forecasts.

  • 2 / 5 — noise

    Helios is already 15-25% faster than Vera Rubin for training and therefore wins the next rack cycle.

    Sources: AMD projections reported at Advancing AI

    Noise as stated. The comparison is forward-looking, workload-sensitive, and not independently reproduced on delivered customer racks. Architecture competition is real; performance victory is not yet evidence.

House measurement

One number we measured ourselves.

Filing-derived: 3 readings, measured by this publication.

Everything else in this issue cites someone’s data. This section is the publication’s own: one data point per week, measured or computed from primary documents, with the method stated so you can check it.

Filing-derived

On a representative 1:4 input-to-output workload, Opus 5's list-price bill is 50% below a Fable 5 rate card at twice the price — before cache or adaptive-thinking effects.

Method: House rate-card calculation using the launch pricing relationship stated by Anthropic. Representative workload: 1 million uncached input tokens and 4 million output tokens. Opus 5 at $5 input and $25 output costs $5 + (4 x $25) = $105. A Fable 5 rate card at approximately twice those rates costs $10 + (4 x $50) = $210. Difference: $105, or 50%. This is a rate-card scenario, not an observed task benchmark.

  • Opus 5 representative bill

    $105

    1M uncached input + 4M output tokens at $5/$25 per million

  • Fable 5 representative bill

    $210

    Same token mix at the approximately 2x premium rate relationship

  • Rate-card compression

    50%

    $105 less on the same token volumes, before behavior differences

Any Fable-heavy production portfolio should run an Opus 5 substitution test before renewal. The savings ceiling is large enough that even partial workload migration matters, but only task-level evaluation can determine the realized amount.

Caveats: Models may consume different token volumes and achieve different completion rates; adaptive thinking, caching, retries, and fast mode change realized cost. The 1:4 token mix is illustrative, and the calculation does not claim equal quality.

Sources: Anthropic Claude Opus 5 launch and pricing

Synthesis

The week, reasoned through.

2 cross-domain connections, 5 hypotheses tested (1 under pressure), 2 patterns tracked.

Reporting says what happened; this section says what it means when you put the pieces together. Every inference is labeled by type, linked to its evidence, and held against the working framework — so when the reasoning is wrong, you can see exactly where.

Connecting the dots

  • Inductive

    81%

    confidence

    Frontier economics are compressing at the model layer while concentrating at the rack layer, making routing and procurement optionality the durable control points.

    1. 01Opus 5 moves near-Fable work into a $5/$25 tier while Gemini Flash lowers high-volume loop cost.
    2. 02Helios makes the accelerator decision a two-rack competition, but each rack still demands roughly a quarter megawatt and deep facility integration.
    3. 03Camellia shows that access to those racks is ultimately bounded by multi-year, customer-funded power infrastructure and curtailment agreements.

    Steel-man

    Model list prices do not guarantee lower task cost, and Helios has not yet proved its performance on delivered racks. The claim survives in narrower form because the available choices expanded even if realized economics remain to be measured.

    Evidence: Anthropic Opus 5 · OpenAI Camellia

  • Deductive

    76%

    confidence

    AI infrastructure is becoming dispatchable utility load, which will force agent and cluster software to treat power availability as runtime state.

    1. 01Camellia commits up to 1 GW of peak curtailment under a 3.2 GW service plan.
    2. 02Helios and Rubin-class racks concentrate 225-246 kW into standardized units whose workloads must checkpoint or move when power is constrained.
    3. 03Metcalfe-style value shifts to the network and orchestration layer that can move jobs across racks, sites, and power windows.

    Steel-man

    Camellia is one unusually large agreement and curtailment may be handled through reserved headroom rather than live workload movement. Even so, a contracted 1 GW interruptible block makes power-aware scheduling an economic requirement somewhere in the system.

    Evidence: OpenAI Camellia curtailment terms · AMD Helios rack coverage

Thesis test

The five standing hypotheses of the working framework, tested deductively against this week’s evidence. A framework that is never strained is not being tested.

  • Hypothesis 1

    supported

    The cycle is accelerating, not slowing.

    Anthropic and Google reset separate price-performance tiers within three days, while AMD moved from accelerator roadmap to full rack and facility design. The release cadence is compressing across software and hardware.

    Against it: Gemini 3.5 Pro remains delayed and Helios production delivery is still ahead, showing that announcement cadence can outrun execution.

    Evidence: Anthropic Opus 5

  • Hypothesis 2

    supported

    Capital is concentrated, returns are diffuse.

    OpenAI's reported $750B plan and $20B Camellia campus concentrate obligation at the frontier, while price compression pushes model savings outward to application builders and users.

    Against it: A successful infrastructure platform could internalize returns through higher utilization and lower unit cost, so diffusion is not guaranteed.

    Evidence: OpenAI Camellia

  • Hypothesis 3

    supported

    Networking is the durable layer.

    Helios depends on merchant Ethernet for both scale-up and 2.4 Tbps-per-GPU scale-out, while curtailment makes cross-rack and cross-site workload movement more valuable. The fabric arbitrates both vendor choice and power availability.

    Against it: No new independent networking revenue print landed this week; the evidence remains architectural and announcement-grade until colo results arrive.

    Evidence: AMD Helios fabric architecture

  • Hypothesis 4

    strained

    Open weights pull the floor up.

    Closed providers moved the price-performance floor this week while Kimi K3 remained hosted-only. The open-weights mechanism cannot claim credit until Moonshot publishes weights and usable license terms.

    Evidence: Moonshot Kimi K3

  • Hypothesis 5

    supported

    Power is the binding constraint for the next 24 months.

    Camellia requires customer-funded infrastructure, staged energization through 2032, and up to 1 GW of curtailment despite extraordinary capital. Helios's 225-245 kW envelope reinforces that rack progress increases facility pressure.

    Against it: The Georgia agreement demonstrates that capital and flexible load can secure a path through the constraint, even if they cannot eliminate the schedule.

    Evidence: OpenAI Camellia power agreement

Pattern watch

  • Inductive4 weeks observed

    Frontier model economics are shifting from list price toward route-aware completed-task cost.

    • W27-W28: closed labs cut headline price and introduced explicit reasoning tiers.
    • W29: Kimi raised the open-lab flagship price while DeepSeek added time-of-day pricing.
    • W30: Opus 5 compressed premium capability and Flash cut both output price and claimed token use.

    Next week: Within two weeks an independent evaluator will publish task cost with token count, latency, and retry data for Opus 5 or Gemini 3.6 Flash.

  • Inductive3 weeks observed

    Power constraints are moving from site-selection inputs into explicit compute operating contracts.

    • W28-W29: secured power, auction scarcity, and permit pauses dominated infrastructure decisions.
    • W30: Camellia contractually pairs 3.2 GW of service with up to 1 GW of peak curtailment.

    Next week: A second large US campus will disclose a material curtailment or dispatchable-load commitment before January 2027.

Second-order effects

  • Trigger: Opus 5 and Gemini Flash compress model economics while adding provider-native routing controls.

    Independent agent platforms lose basic routing as differentiation and must defend on cross-provider policy, observability, evaluation, and workflow-specific verification.

    Horizon: Next 6-12 monthsWho moves: Agent-platform vendors, enterprise AI gateways, model providers, and application engineering teams
  • Trigger: AMD Helios standardizes a 72-GPU open-Ethernet rack with a 225-245 kW envelope.

    Accelerator competition shifts into facilities and financing: power systems, cooling references, software support, and bankable customer offtake become as important as silicon benchmarks.

    Horizon: 2027 deployment cycleWho moves: Neoclouds, electrical and cooling vendors, infrastructure lenders, and large cluster buyers

Strategic outlook

The next twelve months will reward optionality, but not generic multi-vendor slogans. At the model layer, build measured routes: Opus 5 for high-value judgment, Flash for volume, hard stops where fallbacks violate policy, and complete effective-route telemetry. At the rack layer, qualify both Rubin and Helios but tie commitments to delivered thermals, software maturity, and customer workload evidence. At the site layer, assume utilities will demand full infrastructure-cost recovery and dispatchable-load rights for multi-gigawatt campuses. The durable control plane will coordinate all three constraints — model quality, rack availability, and power state — rather than optimize any one in isolation.

Where we differ

Our read against the field.

3 top-tier positions engaged, 2 disagreements on the record.

The best analysts covered this week too. Here is what they said, what we borrow with credit, and where our read genuinely departs from theirs — on the record, so you can score us later.

  • Their take: Opus 5 is primarily another benchmark-leading flagship release.

    Our read: The larger event is economic and operational: near-Fable capability at half the rate, plus fallback and mutable-tool controls that move the API toward an agent control plane.

  • Their take: Helios's projected 15-25% training advantage means AMD has beaten Rubin.

    Our read: AMD has established a credible rack architecture, not a production-performance victory. Delivered systems, thermals, software reliability, and named customer workloads are the adjudicating evidence.

  • Their take: Camellia is mainly another data-center megaproject in OpenAI's spending race.

    Our read: The $20B matters less than the contract: customer-funded infrastructure plus up to 1 GW of utility-directed curtailment is a replicable template for permitting multi-gigawatt load.

Early warning panel

The levers we monitor.

10 metrics tracked — 2 rising, 1 falling, 7 steady.

Current vs prior period. Each metric has a threshold where the read materially changes — this panel flags the inflection before it lands in headlines. Click any metric for the methodology and this-week read.

  • Frontier lab cash position (avg months runway, disclosed-burn labs)

    ~30-40 mo; unchanged cash estimate, but OpenAI's reported infrastructure obligation rises to $750B through 2030vs ~30-40 mo (range; unaudited inputs); no new primary capital — movement was compute sourcing

    Threshold: <18 mo triggers re-rating risk

    What this measures

    No primary financing closed in-window. The new information is commitment intensity: a $20B named site and larger through-2030 plan raise future funding requirements without changing current cash.

  • Hyperscaler capex / AI revenue ratio (top 4 weighted)

    ~5.3-5.8; committed numerator rises with Camellia and OpenAI's reported $750B plan, while segmented AI revenue remains undisclosedvs ~5.0-5.5; Meta Hyperion and Google's Wyoming campus hardened the numerator

    Threshold: >6.0 invites investor pushback at next earnings

    What this measures

    The ratio remains an estimate because AI revenue is not separately reported. Camellia adds hard site economics before revenue catches up, increasing scrutiny on utilization and depreciation.

  • CoreWeave revenue backlog

    $99.4B as of Mar 31; unchanged, with Helios creating a future second-source option rather than a current backlog eventvs $99.4B as of Mar 31 (+284% YoY), restated in Fitch's Jul 16 note

    Threshold: Conversion velocity matters more than gross figure

    What this measures

    No official print in-window. Watch the early-August quarter for conversion and whether customers begin requesting AMD capacity alongside NVIDIA commitments.

  • NVIDIA Q-over-Q data center revenue

    $75.2B Q1 FY27 unchanged; competitive frame tightens as AMD markets a complete 72-GPU Helios rack against Rubinvs $75.2B Q1 FY27; Rubin reported in production but customer-delivery timing remained open

    Threshold: Q2 FY27 guide $91B implies further +21% QoQ

    What this measures

    No NVIDIA earnings event. Helios changes the negotiation set, not current revenue; Aug 26 remains the first hard read on Rubin ramp and competitive response.

  • Open vs closed gap on coding (SWE-Bench / agentic)

    Closed frontier widens at the top with Opus 5; Kimi K3 weights still pending, so the open deployment-control claim remains unresolvedvs ~3 pts on the AA Intelligence Index — K3 at 57.1 vs Fable 5 at 59.9 — weights pending

    Threshold: Sustained open lead reshapes enterprise procurement

    What this measures

    Opus 5 improved the closed tier while K3 remained hosted-only. The score gap matters less than the deployment fact until Moonshot publishes weights and license terms.

  • Sovereign AI commitments (count / aggregate $)

    ~15 / ~$186B; unchanged, while restricted Gemini Cyber access reinforces government-first capability gatingvs ~15 / ~$186B after Japan FRONTia's ¥1T program

    What this measures

    No new sovereign capital commitment qualified in-window. The policy signal is access: specialist cyber capability remains government/partner restricted.

  • PJM 2026/27 capacity auction price ($/MW-day)

    $325.00 unchanged; Camellia shows the emerging workaround — customer-funded grid infrastructure plus contracted peak curtailmentvs $325.00 for 2028/29, at the FERC cap and 6,831 MW short

    Threshold: 11x in 24 months — power is the binding constraint

    What this measures

    No new auction. Camellia does not relax PJM scarcity, but it demonstrates the contract structure large-load utilities may demand in other constrained regions.

  • Time-to-power, busiest US markets (months)

    60-84 unchanged; Camellia's 3.2 GW service is staged across 2028-2032 despite customer-funded infrastructurevs 60-84; New York added a statewide discretionary-permit pause

    What this measures

    Even a flagship, fully funded project carries a multi-year energization schedule. Money can secure a queue position and infrastructure, but it does not collapse construction and permitting time.

  • Cost-per-task, frontier reasoning model

    Premium closed-model floor compresses: Opus 5 at $5/$25 versus Fable 5 at roughly twice the rate; Flash adds lower-token fleet economicsvs AA task floor $0.04 on DeepSeek V4 Pro; named-model median ~$0.63

    Completed-task telemetry must include output tokens, cache hits, fallbacks, and latency tier

    What this measures

    The list-price change is clear, but task cost depends on adaptive thinking and verbosity. Re-run harness evaluations rather than translating rates directly into savings.

  • Custom silicon share of incremental AI compute

    ~34-37% unchanged; Helios strengthens merchant-GPU competition but does not alter the custom-silicon share estimatevs ~34-37%; Google reportedly pitching TPUs into the neocloud channel

    Threshold: >35% materially compresses merchant GPU pricing

    What this measures

    AMD versus NVIDIA is competition within merchant accelerators. The custom-silicon lever moves only when TPU, Trainium, or equivalent deployments change the incremental mix.

Predictions

What we expect next.

4 predictions for the next 30-90 days, confidence 49%-74%.

Each prediction is falsifiable, time-bounded, and tied to a specific signal we will watch. Future issues score these hit, miss, partial, or pending and build a public track record.

Prediction 01

74%

confidence

Software

Artificial Analysis publishes an Opus 5 Intelligence Index result within three points of Claude Fable 5 by August 15, 2026.

Deadline: By August 15, 2026

Trigger: A public Artificial Analysis model page scoring Opus 5 no more than 3.0 Index points below Fable 5 on the then-current methodology.

Prediction 02

68%

confidence

Hardware

At least one named customer reports receiving a production AMD Helios rack for workload qualification by June 30, 2027.

Deadline: By June 30, 2027

Trigger: Customer or AMD announcement naming a delivered 72-GPU MI455X Helios production rack running customer qualification workloads.

Prediction 03

49%

confidence

Power

A second US multi-gigawatt AI campus publicly commits to at least 250 MW of utility-directed peak curtailment by January 31, 2027.

Deadline: By January 31, 2027

Trigger: Utility, developer, or customer filing naming a second campus, curtailment amount of at least 250 MW, and utility dispatch rights.

Prediction 04

66%

confidence

Software

An independent benchmark finds Gemini 3.6 Flash at least 12% cheaper per completed agentic task than Gemini 3.5 Flash by August 31, 2026.

Deadline: By August 31, 2026

Trigger: Independent published harness results comparing total completed-task cost on the same agentic task set and reporting at least 12% savings.

Track record

Scoring prior predictions.

7 prior predictions: 0 hit, 0 miss, 1 partial, 6 pending. Hit rate 0%.

7 predictions across issues so far. Hit rate: 0%. Hits 0, misses 0, partials 1, pending 6.

Prediction 01

72%

confidence

Software

Moonshot publishes Kimi K3 open weights on Hugging Face with a license permitting commercial self-hosting by August 10, 2026.

Deadline: By August 10, 2026

Trigger: Kimi K3 weights live on Hugging Face with published license text; Artificial Analysis reclassifies K3 from proprietary to open-weights.

pendingStill pending as of Jul 25. Moonshot's stated Jul 27 date has not yet arrived; no weights or commercial license were public at publication.

Prediction 02

76%

confidence

Software

DeepSeek retires its legacy deepseek-chat and deepseek-reasoner aliases on July 24 and ships DeepSeek V4 to official GA by July 31, 2026, with peak-hour pricing in effect.

Deadline: By July 31, 2026

Trigger: DeepSeek API docs showing V4 GA model IDs and the alias-retirement notice executed; surge pricing live in the rate card.

partialThe alias-retirement leg hit at 15:59 UTC Jul 24 with hard failure and no redirect. Explicit V4 Pro/Flash IDs and thinking parameter support the migration thesis, but the full GA/pricing trigger is not yet documented strongly enough for a hit.

Prediction 03

68%

confidence

Hardware

SK hynix's July 29 Q2 earnings call discloses that 2027 HBM capacity is substantially sold out or committed under long-term agreements.

Deadline: By July 29, 2026

Trigger: SK hynix Q2 2026 earnings call commentary on 2027 HBM capacity commitments and capex.

pendingThe Jul 29 earnings event is outside this issue's publication date.

Prediction 04

71%

confidence

Networking

The largest colocation operators' Q2 prints show interconnect or fabric revenue growth outpacing overall revenue growth.

Deadline: By August 15, 2026

Trigger: Q2 2026 colo earnings disclosures comparing interconnection or fabric revenue growth with total revenue growth.

pendingThe relevant Q2 prints begin after publication; no qualifying result yet.

Prediction 05

57%

confidence

Power

At least one additional US state announces a statewide restriction on large data-center development by October 31, 2026.

Deadline: By October 31, 2026

Trigger: A governor's order or enacted state legislation pausing or restricting large data-center permitting in a second state.

pendingNo second statewide action qualified this week. Camellia shows a negotiated utility path rather than a moratorium path.

Prediction 06

84%

confidence

Software

DeepSeek V4's official GA pricing does not reset the ultra-cheap floor through August 31, 2026.

Deadline: By August 31, 2026

Trigger: DeepSeek's published API pricing page for GA deepseek-v4-pro keeps off-peak output at or above ¥6 per million tokens.

pendingAlias retirement occurred, but the pricing prediction remains open until the stated deadline and authoritative GA rate card evidence.

Prediction 07

58%

confidence

Software

Gemini 3.5 Pro reaches public general availability with a callable API model ID and published pricing by July 31, 2026.

Deadline: By July 31, 2026

Trigger: Google Gemini API model list and pricing page showing a GA gemini-3.5-pro model ID.

pendingGoogle shipped 3.6 Flash, Flash-Lite, and restricted Flash Cyber instead. Pro remains delayed with six days left on the prediction window.

Track record

The full ledger, misses included.

21 of 70 predictions resolved: 5 hit, 11 partial, 5 miss.

Every prediction this publication has ever made, scored against its own written trigger when the deadline passes — ambiguity resolves against us. Overdue means we haven’t adjudicated yet; it stays visible until we do.

70

predictions made

50%

hit rate (partial = half)

0.139

Brier score (0 = perfect)

0

overdue, unresolved

Calibration by confidence band

  • Bold (<55%)

    No resolved predictions yet — a gap the craft rules now force us to fill.

  • Core (55-80%)

    21 resolved · hit rate 50% vs mean confidence 67%

  • High-conviction (>80%)

    No resolved predictions yet — a gap the craft rules now force us to fill.

Recently resolved

  • partial80% called

    Aggregate 2026 hyperscaler capex revises upward by 10% or more from the $700B baseline.

    Q1 prints (MSFT $190B, GOOG $180-190B, META $125-145B, AMZN $200B reaffirmed) take 2026 aggregate to $695-725B (+77% YoY) vs the $700B W17 baseline. At/near baseline; +10% revision (~$770B) plausible by Q2 print. Score moves to hit if Q2 takes aggregate above $770B.

  • hit66% called

    Samsung's HBM4 supply to NVIDIA is publicly confirmed — via earnings call, company statement, or multi-source supply-chain reporting — by August 31, 2026.

    Hit on the multi-source-reporting trigger: Korean press (Seoul Economic Daily, Korea Herald) reported alongside Samsung's record Q2 guidance that HBM4 — in mass production since February for NVIDIA's Vera Rubin — reached $1B in sales within four months. Caveat: Samsung's Jul 30 divisional results would make it unambiguous from the company itself.

  • hit62% called

    GPT-5.6 reaches broad GA with the Terra tier priced at or below $2.50/$15 per MTok — half of GPT-5.5's rate — confirming a closed-lab repricing cycle rather than a one-off Sonnet 5 cut, by August 31, 2026.

    Hit, seven weeks early. GPT-5.6 went GA Jul 9 with Terra at exactly $2.50/$15 per MTok. Grok 4.5's $2/$6 launch the day before makes it a three-vendor repricing cycle (Sonnet 5, Terra, Grok 4.5), not a one-off.

  • hit66% called

    At least one major enterprise platform ships an admin control specifically for scheduled/background coding or app-building agents by August 31, 2026.

    Hit. GitHub shipped Copilot agent session streaming to public preview (Jul 2) — SIEM/Purview streaming of all agent sessions — on top of its agent control plane, and GitHub also added AI-credit session limits covering background agents (Jul 1, per Agent Techniques coverage).

  • partial65% called

    Broadcom, Marvell, or NVIDIA announces a new CPO/1.6T production design win or revenue guide uplift tied to AI networking before August 31, 2026.

    Arista's 1.6T 7060XE7 portfolio on Broadcom's Tomahawk 6 (Jun 9) is a fresh Broadcom 1.6T production design win, satisfying the 1.6T leg; no co-packaged-optics production win or vendor revenue-guide uplift yet. Tracking to a full hit by deadline.

  • partial60% called

    Expanded Beam Optical MSA publishes a v1.0 spec within 90 days of launch (May 12), with at least one in-production deployment announced by a hyperscaler member (AMD, Cisco, Meta, Oracle).

    EBO MSA membership expanded 17 to 23 vendors May 18 (HPE marquee addition, Bellwether, JPC Connectivity, Mixx, TIME, TFC). v1.0 spec not yet published. Member growth is positive signal but spec + in-production deployment still pending. On track.

Watchlist

On the radar this week.

5 catalysts to watch, starting Jul 27.

Specific catalysts that would change the read materially. Watching these tells us whether the thesis is strengthening or weakening.

  • Jul 27

    Kimi K3 weights and license

    The promised drop decides whether last week's open-frontier thesis becomes a deployment fact or a missed roadmap commitment.

  • Jul 29-30

    SK hynix, Microsoft, Samsung, and colo Q2 prints

    The cluster tests 2027 HBM scarcity, hyperscaler capex absorption, and whether interconnect revenue continues to outgrow the base business.

  • By Jul 31

    Gemini 3.5 Pro and DeepSeek V4 prediction deadlines

    Both standing predictions require public model IDs and pricing evidence; shipped adjacent products do not satisfy their triggers.

  • By Aug 15

    Independent Opus 5 evaluations

    The price compression is factual; independent intelligence, coding, token-use, and cost-per-task results decide how much workload should move.

  • H2 2026

    Helios delivery and facility qualification

    Watch named customer racks, measured thermals, software readiness, and real workloads — the evidence needed to turn a credible architecture into a credible supply alternative.

Companion reads

The rest of the spine.

The AI Stack Weekly is the cross-stack flywheel read. Pair it with the model-and-tree spine and the working framework to get the full picture.

Edits this issue

  • W30-r2 added AMD primary Helios sources, CoreWeave's measured Rubin efficiency result, NVIDIA Spectrum-6, the AMD-Anthropic partnership, Hut 8 Beacon Point Phase 2, and OpenAI's Hugging Face evaluation incident.
  • Added the European Commission's Jul 20 final Article 50 transparency guidance; the covered obligations apply from Aug 2, 2026: https://digital-strategy.ec.europa.eu/en/news/commission-publishes-guidelines-transparency-obligations-providers-and-deployers-certain-ai-systems
  • Added Claude Opus 5 and Gemini 3.6 Flash to the LLM Evolutionary Tree through the Model Pulse tree delta.
  • Prediction p62 moved to partial after DeepSeek executed the hard alias retirement; p57 remains pending because Google shipped Flash variants rather than Gemini 3.5 Pro.
  • No living thesis or market document changed: this week strengthens existing price-compression, networking, and power-constraint hypotheses without creating a durable shockwave that warrants rewriting them.

About this brief

Compiled from public announcements, SEC filings, earnings transcripts, and official lab and vendor publications. Every quantitative claim is graded 1–5 on source quality. Claims graded 2 or below are flagged as noise. The thesis the brief defends is published separately and updated only when a hypothesis materially changes.

Authorship

Written by Brian Letort. Independent analysis. All sources cited are public. Not investment guidance.

Operate. Publish. Teach.