Skip to content

AI Market · Forecast desk · Live · networking

43%

A major model-serving platform or AI gateway publishes per-endpoint accuracy, precision, or output-token-limit disclosures for the open-weight models it serves by January 31, 2027.

Probability history

0%25%50%75%100%08-0809-06deadline

Called at 31% on August 8, 2026. Repriced 1 time; now 43%. 147 days to deadline.

How the number was built

Final method: stacked-fixed-weights (resolved n=57 < 80). Each rung is an independent estimate; the final is a stacked combination.

Stated (issue date)31%
Reference class (n=8)39%
Recalibrated50%
Feature model38%
Council41%
Final43%

Drivers

  • Reference class qtype:timing runs 0.39 (n=8) vs stated 0.31.
  • Hierarchical calibration moves stated 0.31 to 0.50 (lens networking, kappa_stated=24.0, kappa_lens=200.0).
  • Feature model at 0.38: conjuncts=2, named disclosure=False, horizon=176d.
  • Council of 5 at 0.41.

Cautions

  • all 5 votes came from one vendor (openai); extremization disabled, persona diversity only

Council

5 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view, aggregated by log odds mean without extremization.

29%

grid and power economist · openai/gpt-5.6-sol

Price and latency are already natural endpoint-level routing attributes, but accuracy is harder: it requires a disclosed benchmark, reproducible methodology, and endpoint-specific results rather than model-level scores. The four-month window permits a dashboard or documentation update, yet the literal all-three requirement makes a full hit materially less likely than publication of only price and latency. The networking and timing reference classes center near 0.32–0.40, but the strict conjunction and likely accuracy bottleneck justify a discount. The forecaster’s 0.31 is therefore reasonable but slightly high.

43%

semiconductor supply chain analyst · openai/gpt-5.6-sol

Price is commonly documented per model or provider endpoint, and latency is increasingly exposed in routing dashboards, making partial credit plausible within 147 days. The binding leg is accuracy: it is workload-dependent, difficult to standardize, and usually reported per model or benchmark rather than per serving endpoint. Requiring a platform-controlled docs or status-page publication further reduces the full-hit probability. Accounting for a meaningful chance of partial credit puts the expected outcome above the forecaster’s 0.31, but below the overall base rate.

53%

capital markets credit analyst · openai/gpt-5.6-sol

Price and observed latency are already natural endpoint-level routing inputs for major model-serving platforms, making at least partial resolution plausible within 147 days. Accuracy is the difficult leg: it requires repeatable evaluations tied to each provider endpoint, version control, and public documentation rather than model-level benchmarks. The strict per-endpoint wording therefore keeps a full hit materially below even odds. However, the explicit partial-credit policy raises expected resolution value above the forecaster’s 0.31 and above the networking base rate.

41%

superforecaster statistician · openai/gpt-5.6-sol

I use the timing base rate of 0.389, cross-checked against networking at 0.318 and multi-leg questions at 0.433. The existential wording (“a major” platform) and partial credit help: endpoint pricing is common, and latency reporting is increasingly plausible. But a full hit requires the harder accuracy metric at endpoint granularity, published in qualifying platform documentation by a fixed date. That conjunctive disclosure requirement offsets the four-month window. The digest provides no platform-specific evidence. The forecaster’s lone sub-0.50 calibration observation is uninformative, while overconfidence appears mainly in higher bands.

38%

geopolitics regulation analyst · openai/gpt-5.6-sol

The 147-day window and competition among several major serving platforms create multiple chances for a launch, while price and latency reporting are already natural platform features. Accuracy is the limiting leg: it requires benchmark definitions, version control, and comparability, and platforms often publish it by model rather than by endpoint. The strict requirement that all metrics be per endpoint and appear in platform documentation or a status page materially lowers the full-hit probability. Partial-credit treatment offsets some of that downside. The networking, timing, and multi-leg reference classes center around 0.32–0.43, supporting an estimate near their midpoint.

Evidence the engine used

direct · undated

OpenRouter Quickstart Guide

Cited by capital-markets-credit-analyst: Price and observed latency are already natural endpoint-level routing inputs for major model-serving platforms, making at least partial resolution plausible within 147 days. Accuracy is the difficult leg: it requires repeatable evaluations tied to each provid