---
title: >-
  The agent bill moved from tokens to the whole system, just as power and financing became the
  binding constraints
publication: The AI Stack Weekly
slug: 2026-W38
issueNumber: 22
isoYear: 2026
isoWeek: 38
publishedAt: '2026-09-19'
canonicalUrl: https://brianletort.ai/industry/weekly/2026-W38
pdfUrl: https://brianletort.ai/downloads/ai-stack-weekly-2026-W38.pdf
schemaVersion: 2026.05.02
flywheelArc: all-three
capitalFlow:
  - category: Frontier Labs
    capitalIn: No disclosed primary financing in the observation window
    capitalInPrior: Unknown
    capitalInDirection: flat
    revenueOut: Undisclosed
    revenueOutPrior: Unknown
    revenueOutDirection: flat
    burnToRevenue: Unknown
  - category: Hyperscaler-Hosted
    capitalIn: No new category-wide financing disclosed
    capitalInPrior: >-
      Unknown as a category total; two constituents disclosed $28.5B of quarterly capital
      expenditure and an announced plan of at least EUR 13B over two years respectively
    capitalInDirection: flat
    revenueOut: No new AI-segment revenue disclosure
    revenueOutPrior: >-
      Unknown as a category total; one constituent disclosed triple-digit cloud infrastructure
      revenue growth and $664B of contracted backlog
    revenueOutDirection: flat
    burnToRevenue: Unknown
  - category: Neoclouds
    capitalIn: $3.0B convertible notes, with a $500M buyer option
    capitalInPrior: >-
      Unknown as a category total; in-window instruments comprise an up-to-$3.1B undrawn facility,
      A$1.1B of subordinated convertible notes, and $375M firm of an $875M headline
    capitalInDirection: flat
    revenueOut: No new revenue disclosure
    revenueOutPrior: Unknown
    revenueOutDirection: flat
    burnToRevenue: Unknown; financing need is disclosed, operating cash conversion is not
  - category: On-Prem / Hybrid
    capitalIn: No comparable disclosed program in the window
    capitalInPrior: Unknown
    capitalInDirection: flat
    revenueOut: Indirect
    revenueOutPrior: Indirect
    revenueOutDirection: flat
    burnToRevenue: Not applicable
levers:
  - metric: Frontier lab cash runway at current burn
    current: Unknown — no lab disclosed cash, burn, or financing in the window
    prior: Unknown — no constituent disclosed cash, burn, or financing in the window
    direction: flat
    threshold: Below 18 months for any disclosed-burn lab
  - metric: Hyperscaler AI capex to disclosed AI revenue ratio
    current: Unknown — no hyperscaler disclosed AI-segment revenue against capital expenditure
    prior: >-
      Unknown as a top-four ratio — one constituent disclosed $28.5B of quarterly capital
      expenditure and roughly $5B of negative free cash flow, with no AI-segment revenue line
    direction: flat
    threshold: Above 6x sustained for two consecutive quarters
  - metric: CoreWeave contracted revenue backlog
    current: $104.2B as of June 30; unchanged pending the next filing
    prior: $104.2B as of June 30, unchanged — no CoreWeave filing landed in the window
    direction: flat
    threshold: Sequential decline, or conversion below 15% annually
  - metric: NVIDIA quarter-over-quarter data center revenue
    current: $89.0B for Q2 FY27; unchanged pending the next NVIDIA print
    prior: $89.0B for Q2 FY27, unchanged — no NVIDIA print in the window
    direction: flat
    threshold: Two consecutive quarters of sequential decline
  - metric: Open-weight to closed-model capability gap on coding
    current: Not comparable — Gemini Live launched without an open-weight peer on the same voice benchmark
    prior: >-
      Not measurable this week — the reference index changed basis twice in four days, breaking
      comparability with the prior reading
    direction: flat
    threshold: Open weights within 2 Index points of the closed leader
  - metric: Sovereign AI program commitments
    current: Unknown — no qualifying government-funded national compute commitment in the window
    prior: >-
      Unknown — no government-funded national compute programme in the window meets the ledger
      method
    direction: flat
    threshold: Above 20 programs or $250B committed
  - metric: PJM capacity auction clearing price
    current: $325.00 per MW-day for 2028/29; unchanged
    prior: $325.00 per MW-day for 2028/29, unchanged — no auction occurred in the window
    direction: flat
    threshold: An auction clearing below the cap, or a FERC-approved increase in the cap itself
  - metric: Time from interconnection request to energization
    current: Unknown — no comparable queue-duration update was published
    prior: >-
      Unknown as a queue duration — the window's only hard figure is 850 MW delivered within one
      quarter by a single operator, which measures delivery rather than energisation or queue time
    direction: flat
    threshold: Below 48 months in two or more major queues
  - metric: Cost per task, frontier reasoning model
    current: Voice layer now $0.005/min input plus $0.018/min output before reasoning and tool charges
    prior: >-
      $3.26 is the floor of the two leading models and $7.63 the other, per index task on the new
      v4.3 basis; no frontier-tier median is computable because the basis changed twice inside the
      window, and neither figure is commensurate with the prior reading
    direction: flat
    threshold: A frontier-tier reasoning model below $1 per million output tokens
  - metric: Custom silicon share of hyperscaler AI compute
    current: Unknown — no audited hyperscaler compute-mix disclosure
    prior: >-
      Unknown — a multi-generation custom inference agreement was signed but discloses a
      warrant-vesting ceiling rather than units, share, or a service date
    direction: flat
    threshold: Above 45% share with audited hyperscaler mix disclosure
predictions:
  - id: p112-live-api-cost-disclosure-dec31
    lens: software
    confidencePct: 43
    deadline: By December 31, 2026
    text: >-
      At least one major voice-agent provider publishes an end-to-end worked cost example that
      includes voice, reasoning, and tool execution by December 31, 2026.
  - id: p113-power-capped-rfp-dec31
    lens: hardware
    confidencePct: 31
    deadline: By December 31, 2026
    text: >-
      A major server, accelerator, or cloud vendor publishes a customer procurement template using
      tokens per megawatt as an acceptance metric by December 31, 2026.
  - id: p114-coreweave-financing-close-nov30
    lens: capital
    confidencePct: 84
    deadline: By November 30, 2026
    text: >-
      CoreWeave completes at least $3 billion of the convertible note offering announced September
      17 by November 30, 2026.
  - id: p115-koa-ga-mar31
    lens: software
    confidencePct: 67
    deadline: By March 31, 2027
    text: Salesforce makes Koa generally available in at least one US region by March 31, 2027.
  - id: p116-fabric-power-telemetry-mar31
    lens: networking
    confidencePct: 37
    deadline: By March 31, 2027
    text: >-
      A major AI networking vendor adds workload-level token-throughput correlation to a generally
      available fabric telemetry product by March 31, 2027.
predictionsPrior:
  - id: p105-second-operator-delivered-capacity-mar31
    lens: capital
    outcome: pending
    deadline: By March 31, 2027
    text: >-
      A publicly traded operator other than Oracle discloses, for a specific reporting period, both
      a megawatt capacity figure delivered or placed in service and a unit count of AI accelerators
      delivered, by March 31, 2027.
  - id: p106-loviisa-fid-jun30
    lens: power
    outcome: pending
    deadline: By June 30, 2027
    text: >-
      Fortum announces an approved investment decision covering at least EUR 300 million of the EUR
      700 million of Loviisa life-extension capital expenditure currently disclosed as pending, by
      June 30, 2027.
  - id: p107-runtime-manifest-mar31
    lens: software
    outcome: pending
    deadline: By March 31, 2027
    text: >-
      A major model provider or evaluation publisher ships a machine-readable runtime or harness
      manifest that ties a published score to a reproducible configuration, by March 31, 2027.
  - id: p108-pjm-large-load-filing-dec31
    lens: power
    outcome: pending
    deadline: By December 31, 2026
    text: >-
      PJM's Section 205 filing on large computational loads is docketed by December 31, 2026 and
      carries a telemetry or remote-disconnect requirement, not merely a ride-through envelope or a
      ramp-rate limit.
  - id: p109-1600zr-two-vendors-jun30
    lens: networking
    outcome: pending
    deadline: By June 30, 2027
    text: >-
      At least two distinct vendors announce 1600ZR-conformant coherent pluggable optics products by
      June 30, 2027.
  - id: p110-agents-api-residency-mar31
    lens: software
    outcome: pending
    deadline: By March 31, 2027
    text: >-
      OpenAI's managed Agents API supports zero data retention or a non-US data residency option by
      March 31, 2027.
  - id: p111-custom-inference-service-date-mar31
    lens: hardware
    outcome: pending
    deadline: By March 31, 2027
    text: >-
      Qualcomm or its counterparty discloses a named service date, first-deployment date, or unit
      volume for the multi-generation custom AI inference agreement, by March 31, 2027.
signalScores:
  - 4
  - 3
  - 2
  - 2
keyTakeaways:
  - >-
    Voice agents became execution systems: Google now keeps dialogue running while tools and deeper
    reasoning work in parallel.
  - >-
    Infrastructure efficiency moved from component benchmarks to tokens per megawatt, with one
    measured Blackwell deployment lifting throughput 24% inside the same power budget.
  - >-
    Capital markets supplied the counter-signal: CoreWeave sought $3B of convertible debt while debt
    tied to an Oracle-leased campus traded below par.
  - >-
    The operating decision is to price agent work end to end, including the conversational layer,
    reasoning model, tools, recovery, and power.
byTheNumbers:
  - value: $0.005/min
    label: Gemini 3.8 Live audio input
  - value: $0.018/min
    label: Gemini 3.8 Live audio output
  - value: +24%
    label: Measured cluster token throughput
  - value: $3.0B
    label: CoreWeave convertible offering
  - value: 89-91¢
    label: Reported price of Project Jupiter loans
---

# The agent bill moved from tokens to the whole system, just as power and financing became the binding constraints

*Issue 22 · Week 38 of 2026 · Published 2026-09-19*

## Executive summary

- Voice agents became execution systems: Google now keeps dialogue running while tools and deeper reasoning work in parallel.
- Infrastructure efficiency moved from component benchmarks to tokens per megawatt, with one measured Blackwell deployment lifting throughput 24% inside the same power budget.
- Capital markets supplied the counter-signal: CoreWeave sought $3B of convertible debt while debt tied to an Oracle-leased campus traded below par.
- The operating decision is to price agent work end to end, including the conversational layer, reasoning model, tools, recovery, and power.

**By the numbers.**

- **$0.005/min** — Gemini 3.8 Live audio input
- **$0.018/min** — Gemini 3.8 Live audio output
- **+24%** — Measured cluster token throughput (Lambda result reported by NVIDIA)
- **$3.0B** — CoreWeave convertible offering
- **89-91¢** — Reported price of Project Jupiter loans

## Big Story

The week's original connection is not that voice models improved; it is that the conversational layer, reasoning layer, tools, power envelope, and financing structure can now be priced separately. Google released Gemini 3.8 Live at $0.005 per minute of audio input and $0.018 per minute of output while allowing tool calls and deeper reasoning to continue behind an uninterrupted conversation. NVIDIA then reported a measured Blackwell deployment where factory-level power management raised cluster token throughput from about four million to five million tokens per second inside the same power budget. At the same time, CoreWeave launched a $3 billion convertible offering and loans tied to an Oracle-leased AI campus were reported at 89 to 91 cents on the dollar. Buyers should therefore model cost per completed workflow across model, harness, tools, recovery, and infrastructure rather than treating the token rate as the cost of the agent.

Flywheel arc: `all-three`.

## Software lens

- **Sep 14.** GitHub adds efficiency, balance, and intelligence tiers to automatic model routing _([GitHub Changelog](https://github.blog/changelog/2026-09-14-configure-cost-and-quality-in-copilot-auto-model-selection/))_
- **Sep 15.** Google releases Gemini 3.8 Live and Live Extended Thinking through the Live API _([Google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/))_
- **Sep 15.** Salesforce introduces Koa, a CRM reasoning model post-trained from NVIDIA Nemotron 3 Super _([Salesforce, NVIDIA](https://www.salesforce.com/news/press-releases/2026/09/15/koa-reasoning-model/))_
- **Sep 18.** GitHub agents gain local Dev Containers and direct pull-request creation from agent sessions _([GitHub Changelog](https://github.blog/changelog/2026-09-18-github-copilot-weekly-releases-september-14/))_

**What this means.** Architects should specify routing policy and the complete metering chain, not a favorite model. Automatic model selection now exposes an explicit cost-quality-latency policy, while live voice can keep a conversation open as tools and reasoning continue; see The Model Pulse for the model and pricing detail.

## Hardware lens

- **Sep 15.** NVIDIA reports Vera Rubin in production and publishes agentic throughput-per-megawatt claims _([NVIDIA](https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/))_
- **Sep 15.** Lambda measures 24% more cluster token throughput under the same Blackwell power budget _([NVIDIA, Lambda](https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/))_
- **Sep 17.** AMD argues for movable workloads across CPUs, accelerators, edge devices, and open interconnects _([AMD](https://newsroom.amd.com/news/building-infrastructure-ai-world/))_

**What this means.** Operators should make tokens per megawatt and performance under a site power cap acceptance metrics for new infrastructure. The measured Lambda result is stronger evidence than NVIDIA's forward Rubin multipliers, but both point to system-level power management becoming as important as accelerator peak throughput.

## Networking lens

- **Sep 15.** NVIDIA frames NVLink, Spectrum-X, ConnectX, BlueField, storage, and cooling as one power-managed factory _([NVIDIA](https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/))_
- **Sep 15.** Cisco highlights cross-domain network, security, and observability context for agentic operations _([Cisco](https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2026/m09/product-keynote-replay-the-platform-for-your-agentic-enterprise.html))_
- **Sep 17.** AMD makes open interconnects and workload portability central to its AI infrastructure posture _([AMD](https://newsroom.amd.com/news/building-infrastructure-ai-world/))_

**What this means.** Network buyers should test the fabric as part of a power-capped system rather than procure switching on headline bandwidth alone. The week's common design point is cross-layer control: compute, memory, fabric, cooling, and observability must expose enough telemetry to optimize one completed-work metric.

## Capital flow

| Category | Capital in | Revenue out | Burn:Revenue | Movement |
|---|---|---|---|---|
| Frontier Labs (OpenAI, Anthropic, Google DeepMind, DeepSeek) | No disclosed primary financing in the observation window (was Unknown, flat) | Undisclosed (was Unknown, flat) | Unknown | Flat on disclosed financing; Anthropic's possible model release and IPO timing remain reported deliberations, not transactions |
| Hyperscaler-Hosted (Azure-OpenAI, AWS-Anthropic, Google Cloud-Gemini, Oracle-OCI) | No new category-wide financing disclosed (was Unknown as a category total; two constituents disclosed $28.5B of quarterly capital expenditure and an announced plan of at least EUR 13B over two years respectively, flat) | No new AI-segment revenue disclosure (was Unknown as a category total; one constituent disclosed triple-digit cloud infrastructure revenue growth and $664B of contracted backlog, flat) | Unknown | Down in credit quality for one project: $18B of Oracle-leased campus loans were reported at 89-91 cents |
| Neoclouds (CoreWeave, Nscale, Crusoe, Lambda, IREN, Zankore, NEXTDC) | $3.0B convertible notes, with a $500M buyer option (was Unknown as a category total; in-window instruments comprise an up-to-$3.1B undrawn facility, A$1.1B of subordinated convertible notes, and $375M firm of an $875M headline, flat) | No new revenue disclosure (was Unknown, flat) | Unknown; financing need is disclosed, operating cash conversion is not | Up sharply as CoreWeave returned to debt and opened an at-the-market equity program |
| On-Prem / Hybrid (Enterprise GPU clusters, sovereign and national programs, open-weight and on-device deployment) | No comparable disclosed program in the window (was Unknown, flat) | Indirect (was Indirect, flat) | Not applicable | Flat on disclosed capital; up in architectural emphasis as vendors stressed portability and private deployment |

### Frontier Labs — detail
No frontier lab disclosed a financing or segment revenue figure during the window. Reuters reported investor scrutiny and possible release timing at Anthropic, but those are market signals rather than ledger inputs.
**Transactions:**
  - **Sep 19.** No qualifying disclosed transaction _([Reuters](https://www.reuters.com/business/anthropic-considers-releasing-new-ai-model-ahead-ipo-sources-say-2026-09-19/))_

### Hyperscaler-Hosted — detail
The week's category signal came from secondary trading rather than a new hyperscaler capital budget. Loans tied to an Oracle-leased New Mexico campus reportedly traded below par amid construction, environmental, and credit concerns.
**Transactions:**
  - **Sep 18.** Project Jupiter loans quoted below par — $18B outstanding _([Reuters citing Financial Times](https://www.reuters.com/business/finance/oracles-18-billion-data-center-debt-under-pressure-ft-reports-2026-09-18/))_

### Neoclouds — detail
CoreWeave's financing is the category's fact home this week. The company launched $3 billion of convertible notes, an option for another $500 million, and an at-the-market program for up to 35 million shares, showing that capacity expansion remains capital intensive even with contracted demand.
**Transactions:**
  - **Sep 17.** CoreWeave convertible debt offering — $3.0B plus $500M option _([CoreWeave, Reuters](https://www.reuters.com/legal/transactional/coreweave-launches-3-billion-convertible-debt-sale-2026-09-17/))_

### On-Prem / Hybrid — detail
No government or enterprise on-premises program supplied a comparable financing number. AMD's portability argument and Salesforce's plan for controlled model deployment are product and architecture signals, not capital commitments.
**Transactions:**
  - **Sep 17.** No qualifying disclosed transaction _([AMD](https://newsroom.amd.com/news/building-infrastructure-ai-world/))_

## Signal vs noise

- **Score 4/5 —** Lambda raised cluster token throughput 24% inside the same power budget by running 19 nodes where 16 full-power nodes were normally allocated.
  - _Sources:_ NVIDIA report of Lambda measurement
  - _Read:_ This is the strongest operating evidence in the window because it names the before and after configuration. Buyers should request the same power-capped test on their workload.
- **Score 3/5 —** Google's Gemini 3.8 Live can continue a conversation while tools and deeper reasoning run in the background.
  - _Sources:_ Google launch posts
  - _Read:_ The capability is shipped through the Live API, but outcome quality and interruption behavior still need workload-specific testing.
- **Score 2/5 —** Vera Rubin delivers up to 30 times more agentic throughput per megawatt than GB300 NVL72.
  - _Sources:_ NVIDIA vendor benchmark using AgentX
  - _Read:_ The number is vendor-reported and workload-specific. Treat it as a test target, not a capacity-plan input, until independently reproduced.
- **Score 2/5 —** Koa produces three times fewer errors than leading models on CRM actions.
  - _Sources:_ Salesforce CRM benchmark
  - _Read:_ The benchmark, dataset, and comparison set are vendor-controlled. Pilot availability is real; the multiple is not yet procurement-grade evidence.


## Levers

| Metric | Current | Prior | Direction | Threshold |
|---|---|---|---|---|
| Frontier lab cash runway at current burn | Unknown — no lab disclosed cash, burn, or financing in the window | Unknown — no constituent disclosed cash, burn, or financing in the window | flat | Below 18 months for any disclosed-burn lab |
| Hyperscaler AI capex to disclosed AI revenue ratio | Unknown — no hyperscaler disclosed AI-segment revenue against capital expenditure | Unknown as a top-four ratio — one constituent disclosed $28.5B of quarterly capital expenditure and roughly $5B of negative free cash flow, with no AI-segment revenue line | flat | Above 6x sustained for two consecutive quarters |
| CoreWeave contracted revenue backlog | $104.2B as of June 30; unchanged pending the next filing | $104.2B as of June 30, unchanged — no CoreWeave filing landed in the window | flat | Sequential decline, or conversion below 15% annually |
| NVIDIA quarter-over-quarter data center revenue | $89.0B for Q2 FY27; unchanged pending the next NVIDIA print | $89.0B for Q2 FY27, unchanged — no NVIDIA print in the window | flat | Two consecutive quarters of sequential decline |
| Open-weight to closed-model capability gap on coding | Not comparable — Gemini Live launched without an open-weight peer on the same voice benchmark | Not measurable this week — the reference index changed basis twice in four days, breaking comparability with the prior reading | flat | Open weights within 2 Index points of the closed leader |
| Sovereign AI program commitments | Unknown — no qualifying government-funded national compute commitment in the window | Unknown — no government-funded national compute programme in the window meets the ledger method | flat | Above 20 programs or $250B committed |
| PJM capacity auction clearing price | $325.00 per MW-day for 2028/29; unchanged | $325.00 per MW-day for 2028/29, unchanged — no auction occurred in the window | flat | An auction clearing below the cap, or a FERC-approved increase in the cap itself |
| Time from interconnection request to energization | Unknown — no comparable queue-duration update was published | Unknown as a queue duration — the window's only hard figure is 850 MW delivered within one quarter by a single operator, which measures delivery rather than energisation or queue time | flat | Below 48 months in two or more major queues |
| Cost per task, frontier reasoning model | Voice layer now $0.005/min input plus $0.018/min output before reasoning and tool charges | $3.26 is the floor of the two leading models and $7.63 the other, per index task on the new v4.3 basis; no frontier-tier median is computable because the basis changed twice inside the window, and neither figure is commensurate with the prior reading | flat | A frontier-tier reasoning model below $1 per million output tokens |
| Custom silicon share of hyperscaler AI compute | Unknown — no audited hyperscaler compute-mix disclosure | Unknown — a multi-generation custom inference agreement was signed but discloses a warrant-vesting ceiling rather than units, share, or a service date | flat | Above 45% share with audited hyperscaler mix disclosure |
**Lever detail:**
- **Frontier lab cash runway at current burn.** The required inputs remain unaudited and absent. GPT-Live-1 list pricing and the Agents API's managed-service terms are product and commercial signals, not financing inputs, and no lab published a raise or a revenue figure this week.
- **Hyperscaler AI capex to disclosed AI revenue ratio.** Oracle's filing supplies a real capital-expenditure number and a real backlog number but no AI-segment revenue, which is the denominator the method requires. No hyperscaler disclosed AI-segment operating margin this week, so the ratio stays unpublished rather than estimated.
- **CoreWeave contracted revenue backlog.** Backlog remains a filed stock value awaiting the next quarter. The week's neocloud financings, an undrawn term loan and a convertible note with an investor put, are category structure signals and cannot be substituted into this series.
- **NVIDIA quarter-over-quarter data center revenue.** The in-window NVIDIA news is an Australian capacity aggregation across eight operators and a third-party accelerator joining NVLink Fusion. Neither is a revenue disclosure, and TSMC's record month is a supplier datapoint that cannot be substituted into this series.
- **Open-weight to closed-model capability gap on coding.** This is a measurement failure rather than a capability result. Two open checkpoints shipped in the window, both explicitly cheaper and smaller rather than stronger, and one carries vendor-reported benchmarks with no independent evaluation. See Model Pulse for the lineage read.
- **Sovereign AI program commitments.** The method counts government-funded national compute programmes only. This week's two candidates fail it for different reasons: an at-least-EUR-13B corporate investment plan with a power purchase agreement is private capital, and an up-to-two-gigawatt national target is an aggregation of eight commercial operators' pipelines.
- **PJM capacity auction clearing price.** The prior threshold, a second consecutive auction clearing at the cap, is already satisfied three times over and can no longer move: 2026/27 cleared at $329.17, 2027/28 at $333.44 and 2028/29 at $325.00, each at the approved cap. PJM's own no-cap-or-floor simulation for 2028/29 reports $554.72 per MW-day for the RTO and $776.69 for ComEd, versus the $325 capped result, with $29.7 billion of simulated cleared value against $16.4 billion actual. The simulation does not isolate the cap as the sole cause, but it quantifies how much higher PJM's model clears without it. The in-window development is regulatory rather than price-setting: proposed large computational load ride-through, ramp-rate, telemetry and remote-disconnect requirements were presented on September 8, targeting a federal filing in November 2026. If mandatory curtailable status attaches to gigawatt-scale loads, it changes what capacity payments are actually buying.
- **Time from interconnection request to energization.** No comparable queue-duration update was published. A Texas docket fight over forfeiture terms on a 75-megawatt interconnection rule is evidence of queue scarcity and of the cost of holding a position in one, but it is not a duration measurement.
- **Cost per task, frontier reasoning model.** The cost spread is the most durable finding to survive the index revision and is a genuine procurement fact: 57% lower cost for the same rounded composite score. It is reported here as two point observations rather than as movement, because the composite they price changed underneath them.
- **Custom silicon share of hyperscaler AI compute.** The week strengthens the directional case for custom inference silicon on two counts, a multi-generation agreement and a third-party accelerator entering the dominant scale-up fabric, while supplying no unit volumes, deployed utilisation, or mix disclosure from which a share could be computed.

## Predictions

- **`p112-live-api-cost-disclosure-dec31` _[software]_ — At least one major voice-agent provider publishes an end-to-end worked cost example that includes voice, reasoning, and tool execution by December 31, 2026.**
  - Confidence: 43%. Deadline: By December 31, 2026.
  - Trigger: Hit only if a first-party pricing or documentation page shows all three cost components in one worked workflow.
- **`p113-power-capped-rfp-dec31` _[hardware]_ — A major server, accelerator, or cloud vendor publishes a customer procurement template using tokens per megawatt as an acceptance metric by December 31, 2026.**
  - Confidence: 31%. Deadline: By December 31, 2026.
  - Trigger: Hit only with a public first-party RFP, reference architecture, or acceptance guide naming tokens per megawatt.
- **`p114-coreweave-financing-close-nov30` _[capital]_ — CoreWeave completes at least $3 billion of the convertible note offering announced September 17 by November 30, 2026.**
  - Confidence: 84%. Deadline: By November 30, 2026.
  - Trigger: Hit only if a CoreWeave filing states gross proceeds of at least $3 billion from the announced notes.
- **`p115-koa-ga-mar31` _[software]_ — Salesforce makes Koa generally available in at least one US region by March 31, 2027.**
  - Confidence: 67%. Deadline: By March 31, 2027.
  - Trigger: Hit only if Salesforce documentation marks Koa generally available, not pilot or beta, in a named US region.
- **`p116-fabric-power-telemetry-mar31` _[networking]_ — A major AI networking vendor adds workload-level token-throughput correlation to a generally available fabric telemetry product by March 31, 2027.**
  - Confidence: 37%. Deadline: By March 31, 2027.
  - Trigger: Hit only if public product documentation joins network telemetry to model token throughput for a named workload.

### Prior predictions scored

- `p105-second-operator-delivered-capacity-mar31` _[capital]_ — **PENDING** — A publicly traded operator other than Oracle discloses, for a specific reporting period, both a megawatt capacity figure delivered or placed in service and a unit count of AI accelerators delivered, by March 31, 2027.
- `p106-loviisa-fid-jun30` _[power]_ — **PENDING** — Fortum announces an approved investment decision covering at least EUR 300 million of the EUR 700 million of Loviisa life-extension capital expenditure currently disclosed as pending, by June 30, 2027.
- `p107-runtime-manifest-mar31` _[software]_ — **PENDING** — A major model provider or evaluation publisher ships a machine-readable runtime or harness manifest that ties a published score to a reproducible configuration, by March 31, 2027.
- `p108-pjm-large-load-filing-dec31` _[power]_ — **PENDING** — PJM's Section 205 filing on large computational loads is docketed by December 31, 2026 and carries a telemetry or remote-disconnect requirement, not merely a ride-through envelope or a ramp-rate limit.
- `p109-1600zr-two-vendors-jun30` _[networking]_ — **PENDING** — At least two distinct vendors announce 1600ZR-conformant coherent pluggable optics products by June 30, 2027.
- `p110-agents-api-residency-mar31` _[software]_ — **PENDING** — OpenAI's managed Agents API supports zero data retention or a non-US data residency option by March 31, 2027.
- `p111-custom-inference-service-date-mar31` _[hardware]_ — **PENDING** — Qualcomm or its counterparty discloses a named service date, first-deployment date, or unit volume for the multi-generation custom AI inference agreement, by March 31, 2027.

## Synthesis

### Connecting the dots

- **AI procurement is moving from model selection to system-yield contracting.** _[abductive, 76% confidence]_
  1. Google split live-agent audio into input and output meters while background reasoning and tools continue.
  2. GitHub exposed cost, quality, and latency as selectable routing policy.
  3. NVIDIA and Lambda measured completed token throughput under a fixed power budget.
  Steel-man: These are vendor-defined meters and one customer measurement, not a standardized contracting unit. The claim is bounded to procurement direction, not current market practice.
  Evidence: [Google Gemini 3.8 Live developer pricing](https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio/), [GitHub auto model tiers](https://github.blog/changelog/2026-09-14-configure-cost-and-quality-in-copilot-auto-model-selection/), [NVIDIA and Lambda power-capped measurement](https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/)
- **Infrastructure credit risk is becoming an operational architecture input rather than a finance-only concern.** _[inductive, 69% confidence]_
  1. CoreWeave sought $3 billion of convertible debt plus an equity-sale facility to fund operations.
  2. Loans tied to an Oracle-leased campus were reported below par amid project-specific concerns.
  3. Hardware vendors simultaneously shifted their headline metric to useful work inside a fixed megawatt.
  Steel-man: Both financings may be issuer- or project-specific and demand remains strong. The inference is about diligence scope, not a sector-wide credit event.
  Evidence: [CoreWeave financing](https://www.reuters.com/legal/transactional/coreweave-launches-3-billion-convertible-debt-sale-2026-09-17/), [Oracle-leased campus debt](https://www.reuters.com/business/finance/oracles-18-billion-data-center-debt-under-pressure-ft-reports-2026-09-18/)

### Thesis test

- **Hypothesis 1 — Software demand pulls hardware and network investment forward.** — **SUPPORTED**. Parallel voice reasoning increases the number of simultaneous model, tool, and memory paths, while infrastructure vendors answer with system-level throughput optimization. Against it: No buyer disclosed incremental capacity ordered specifically for Gemini Live. Evidence: Gemini Live launch, NVIDIA AI factory result
- **Hypothesis 2 — Hardware efficiency expands economically viable AI demand.** — **SUPPORTED**. Lambda's measured 24% throughput gain inside the same power budget directly lowers the infrastructure consumed per token. Against it: One Blackwell deployment does not establish portability across models or facilities. Evidence: Lambda DSX measurement
- **Hypothesis 3 — Networking becomes a first-order limiter as AI systems scale.** — **SUPPORTED**. The week's vendor architectures treat fabric, memory, storage, and cooling as one controlled system rather than separable components. Against it: No in-window outage or benchmark isolated networking as the binding bottleneck. Evidence: NVIDIA factory architecture, AMD open infrastructure posture
- **Hypothesis 4 — Capital follows visible utilization and contracted demand.** — **STRAINED**. CoreWeave accessed large financing, but below-par project debt shows that contracted demand no longer neutralizes construction, environmental, and concentration risk. Evidence: CoreWeave notes, Project Jupiter loans
- **Hypothesis 5 — Open ecosystems gain when switching costs become material.** — **UNTESTED**. AMD emphasized open interconnects, but the window supplied no comparable buyer migration or audited share shift. Evidence: AMD infrastructure essay

### Pattern watch

- **Agent cost decomposes into independently managed system meters** _[inductive, 2 weeks observed]_
  - W37: GPT-Live-1 split voice from separately billed reasoning.
  - W38: Gemini Live split audio input and output while tools run in parallel.
  Next week: A major provider will publish a worked multi-meter agent cost example before year end.
- **Power-capped throughput replaces peak component performance** _[inductive, 2 weeks observed]_
  - W37: long-context releases led on cache bytes per token.
  - W38: Lambda reported tokens per second and performance per watt inside a fixed power budget.
  Next week: At least one infrastructure launch next week will headline useful work per watt or megawatt rather than peak FLOPS.

### Second-order effects

- **Trigger:** Voice agents continue talking while background tools and reasoning execute. **Effect:** FinOps must attribute one user interaction across several concurrent meters and failure domains. _(Horizon: Next two quarters. Who moves: CIOs, application architects, and AI platform teams.)_
- **Trigger:** AI campus debt trades below par while new capacity still seeks billions in financing. **Effect:** Workload portability and staged capacity commitments become credit-risk mitigations, not only architecture preferences. _(Horizon: Next 12 months. Who moves: Cloud buyers, lenders, neoclouds, and infrastructure vendors.)_

### Strategic outlook

Over the next 12 months, the winning capital posture is optionality with measurement: buy capacity in stages, require power-capped workload tests, preserve model and fabric portability where practical, and price agents as complete workflows rather than token endpoints. The risk is no longer only overpaying for a model that becomes obsolete. It is locking a workflow to a conversational meter, a reasoning meter, a tool chain, and an infrastructure financing structure that can each move independently.


## Track record

Cumulative ledger: **116 predictions made**, 57 resolved (23 hit / 15 partial / 19 miss), 59 pending, 0 overdue. Hit rate (partial = half): **54%**. Brier score: **0.200** (0 = perfect, 0.25 = coin-flip).

Calibration by confidence band:
- Bold (<55%): 1 resolved, hit rate 100% vs mean confidence 43%
- Core (55-80%): 55 resolved, hit rate 52% vs mean confidence 66%
- High-conviction (>80%): 1 resolved, hit rate 100% vs mean confidence 84%

Recently resolved:
- **HIT** (called at 72%): NVIDIA files exhibits with the 10-Q for the quarter ended July 26, 2026 that translate the SB Energy PORTS-Pike residual-value guaranty into a per-quarter contingent-obligation disclosure and identify the OpenAI affiliate as tenant, by October 31, 2026. — NVIDIA filed the Form 10-Q for the quarter ended July 26, 2026 on August 26, 2026 — inside the window. It satisfies all three trigger elements: guarantees 'capped at a total of $105 billion' with an exposure table of $3.5B AI-cloud guarantees plus $105.0B SB Energy for $108.5B total; effectiveness conditioned on SB Energy satisfying applicable ready-for-service conditions as each of nine phases is placed in service from fiscal 2029; and the tenant identified as 'an affiliate of OpenAI Group PBC' at the PORTS Technology Campus in Pike County, Ohio. Exhibit 10.1 is the Form of Residual Value Guaranty.
- **PARTIAL** (called at 80%): Aggregate 2026 hyperscaler capex revises upward by 10% or more from the $700B baseline. — Q1 prints (MSFT $190B, GOOG $180-190B, META $125-145B, AMZN $200B reaffirmed) take 2026 aggregate to $695-725B (+77% YoY) vs the $700B W17 baseline. At/near baseline; +10% revision (~$770B) plausible by Q2 print. Score moves to hit if Q2 takes aggregate above $770B.
- **HIT** (called at 43%): Z.ai publishes GLM-5.3 weights to Hugging Face by September 15, 2026, closing the two-week window promised at the model's August 14 announcement. — Z.ai published the full 753B-parameter GLM-5.3 weights to Hugging Face at zai-org/GLM-5.3 on August 27–28, 2026 — in-window and inside the trigger's September 15 window, distinct from GLM-5.2 — after GLM-5.3-Flash MIT weights landed Aug 26. The material nuance is licensing, not availability: GLM-5.3 ships under a bespoke GLM-5.3 license rather than MIT, requiring Z.AI security review before commercial use by any Model-as-a-Service operator whose group revenue exceeds $10B over any 12 consecutive months.
- **HIT** (called at 66%): An independent benchmark finds Gemini 3.6 Flash at least 12% cheaper per completed agentic task than Gemini 3.5 Flash by August 31, 2026. — Artificial Analysis measured Gemini 3.6 Flash at $0.50 average cost per completed agentic task versus $0.59 for 3.5 Flash — a 15% reduction, above the 12% cheaper-per-task bar — before Aug 31.
- **HIT** (called at 84%): DeepSeek V4's official GA pricing does not reset the ultra-cheap floor: off-peak deepseek-v4-pro output pricing stays at or above ¥6 (~$0.85) per MTok through August 31, 2026 — the kill-condition test for this issue's price-band-convergence claim. — DeepSeek's official API pricing page kept GA deepseek-v4-pro off-peak output at $1.98/MTok (~¥14+) through Aug 31 — well above the ¥6 (~$0.85)/MTok ultra-cheap floor the trigger set as the kill condition.
- **HIT** (called at 64%): At least one major agent platform (OpenAI, Anthropic, GitHub, or Cursor) ships product-level per-task or per-harness cost telemetry or routing controls — beyond session budget caps — by August 31, 2026. — Cursor shipped Cursor Router in July 2026 with Auto Balance/Intelligence routing controls and published measured cost-per-commit figures ($4.63–$6.76) from live traffic — product-level harness routing and cost telemetry beyond session budget caps.

## Watchlist

- **Sep 21-25 — Independent Gemini 3.8 Live latency and interruption tests.** The shipped rate card is useful only when paired with measured turn latency, tool-call delay, and interruption recovery.
- **Sep 21-30 — CoreWeave convertible pricing and closing.** Coupon, conversion premium, hedging cost, and final proceeds will show what capital markets charge for the next unit of neocloud capacity.
- **Sep 21-Oct 2 — Project Jupiter credit and permitting response.** Any lender, county, or Oracle disclosure could distinguish temporary syndication pressure from a material construction risk.
- **Oct 2026 — Koa pilot evidence and open-beta timing.** Named pilots need workload definitions and error baselines before the vendor's three-times-fewer-errors claim can inform procurement.

## Changelog

- Authored from verified public sources dated September 14-19, 2026, with primary vendor material used for product claims and Reuters used for financing and market reports.
- No matured W37 prediction deadline fell inside the observation window; all seven prior predictions remain pending.
- The fact home for CoreWeave financing is Capital Flow; other publications use only short cross-references.

---

Source of truth: `src/data/industry/weekly/2026-W38.ts`. Canonical HTML: <https://brianletort.ai/industry/weekly/2026-W38>. PDF: <https://brianletort.ai/downloads/ai-stack-weekly-2026-W38.pdf>.
