Skip to content

AI Market · Forecast desk · Live · software

25%

The highest single ISO week of OpenRouter aggregate token volume in September 2026 exceeds the Ox Alpha stealth-week peak (week of August 20–26, 2026) by at least 15%, by September 30, 2026.

Probability history

0%25%50%75%100%08-2909-06deadline

Called at 40% on August 29, 2026. Repriced 1 time; now 25%. 24 days to deadline.

How the number was built

Final method: stacked-fixed-weights (resolved n=57 < 80). Each rung is an independent estimate; the final is a stacked combination.

Stated (issue date)40%
Reference class (n=19)65%
Recalibrated51%
Feature model34%
Council12%
Final25%

Drivers

  • Reference class qtype:threshold runs 0.65 (n=19) vs stated 0.40.
  • Hierarchical calibration moves stated 0.40 to 0.51 (lens software, kappa_stated=24.0, kappa_lens=200.0).
  • Feature model at 0.34: conjuncts=1, named disclosure=False, horizon=32d.
  • Council of 5 at 0.12 with dissent.

Cautions

  • all 5 votes came from one vendor (openai); extremization disabled, persona diversity only

Council

5 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view, aggregated by log odds mean without extremization. Dissent: geopolitics-regulation-analyst at 0.34 vs capital-markets-credit-analyst at 0.04: The numeric bar and current weekly volume are absent, so there is no observable cushion to extrapolate. / The supplied written trigger does not state a numeric bar, so the required comparison cannot be made from the question.

31%

grid and power economist · openai/gpt-5.6-sol

No numeric bar, OpenRouter baseline, or September-to-date volume is supplied, so the written trigger cannot be independently tested. Structurally, taking the maximum across several ISO weeks makes staying below a bar less likely than a single-week comparison. Late-August and early-September model releases could also produce traffic spikes, though release cadence is only an indirect signal. The 2026-09-30 vintage rule is clear, but treatment of the ISO week spanning September 28–October 4 may be ambiguous; the stated policy resolves ambiguity against the forecaster. I therefore put the event modestly below the forecaster’s 0.4.

8%

semiconductor supply chain analyst · openai/gpt-5.6-sol

No numeric token-volume bar appears in the supplied written trigger, and the digest provides no OpenRouter weekly volume observations. The resolution rule requires comparison with that written trigger, while ambiguity explicitly resolves against the forecaster. Therefore the affirmative “stays below” outcome is unlikely to score without a recoverable, unambiguous threshold, irrespective of underlying usage trends. The forecaster’s ledger offers little help: only one prior forecast was below 0.50 confidence.

4%

capital markets credit analyst · openai/gpt-5.6-sol

The supplied written trigger does not state a numeric bar, so the required comparison cannot be made from the question. It also leaves unclear whether an ISO week merely overlapping September qualifies and how the incomplete week beginning September 28 is treated at the September 30 vintage. Because the rule requires scoring the written trigger literally and ambiguity resolves against the forecaster, these omissions make a positive resolution unlikely regardless of the underlying volume trend.

6%

superforecaster statistician · openai/gpt-5.6-sol

I use the software-ledger base rate of 0.452; the broader threshold base rate is 0.650 but is less specific. The numerical bar is absent from the supplied written trigger, and the digest provides neither OpenRouter’s current volume nor its recent trend. Under literal scoring and an ambiguity policy that resolves against the forecaster, that omission is a major failure risk. The claim also concerns the maximum across all September ISO weeks, so one spike can defeat it. The forecaster’s low-confidence calibration bin has only one observation and provides no useful correction.

34%

geopolitics regulation analyst · openai/gpt-5.6-sol

The numeric bar and current weekly volume are absent, so there is no observable cushion to extrapolate. Literal resolution and ambiguity against the forecaster materially reduce the chance of success, especially because September 30 cuts through an ISO week and historical revisions use that day’s vintage. Several late-August and early-September model releases could also generate a weekly spike. The software reference rate of 0.452 is more relevant than the stronger generic threshold rate; after the specification and boundary-week penalties, I place this below the forecaster’s 0.4.

Evidence the engine used

direct · undated

LLM Rankings | OpenRouter

Cited by geopolitics-regulation-analyst: The numeric bar and current weekly volume are absent, so there is no observable cushion to extrapolate. Literal resolution and ambiguity against the forecaster materially reduce the chance of success, especially because September 30 cuts through an ISO week a