Skip to content

AI Market · Forecast desk · Live · software

54%

Artificial Analysis publishes Endpoint Accuracy Index results covering at least two models beyond the initial GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro set by October 31, 2026.

Probability history

0%25%50%75%100%08-0809-06deadline

Called at 83% on August 8, 2026. Repriced 1 time; now 54%. 55 days to deadline.

How the number was built

Final method: stacked-fixed-weights (resolved n=57 < 80). Each rung is an independent estimate; the final is a stacked combination.

Stated (issue date)83%
Reference class (n=19)65%
Recalibrated56%
Feature model51%
Council55%
Final54%

Drivers

  • Reference class qtype:threshold runs 0.65 (n=19) vs stated 0.83.
  • Hierarchical calibration moves stated 0.83 to 0.56 (lens software, kappa_stated=24.0, kappa_lens=200.0).
  • Feature model at 0.51: conjuncts=1, named disclosure=False, horizon=84d.
  • Council of 5 at 0.55.

Cautions

  • all 5 votes came from one vendor (openai); extremization disabled, persona diversity only

Council

5 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view, aggregated by log odds mean without extremization.

58%

grid and power economist · openai/gpt-5.6-sol

A 55-day window is sufficient for an established benchmark publisher to add and publish endpoint evaluations, but the resolution standard requires the complete stated expansion; partial pages explicitly miss. No digest evidence confirms that the expanded coverage is already live, in testing, or scheduled. Relevant ledger rates cluster around 0.45 for software and 0.62–0.65 for this horizon and threshold structure. The forecaster’s 0.83 confidence band has realized only 0.50, albeit on four cases. I therefore place this near the middle of those reference classes rather than at 0.83.

58%

semiconductor supply chain analyst · openai/gpt-5.6-sol

A 55-day window is sufficient for a software-index expansion, and Artificial Analysis can publish without external hardware or OEM dependencies. However, the supplied digest gives no direct evidence of implementation progress, preview results, or a dated commitment. The resolution rule requires complete coverage; partial pages explicitly miss, increasing execution and scope risk. Relevant ledger rates cluster around 0.45 for software and 0.62–0.65 for this horizon/threshold type. The forecaster’s 0.8–0.9 band has realized only 0.50, so 0.83 appears too aggressive.

57%

capital markets credit analyst · openai/gpt-5.6-sol

Artificial Analysis regularly updates public model and provider benchmarks, and continued model-release activity supports demand for broader endpoint coverage. However, the digest gives no direct evidence that the specific expansion is already live, announced, or scheduled. The exact required expansion is also not reproduced here; because partial pages miss and ambiguity resolves against the forecaster, an ordinary incremental update may not qualify. With only 55 days remaining, I place this near the relevant threshold/horizon base rates but well below the forecaster’s 0.83, especially given weak historical outcomes in its 0.8–0.9 confidence band.

45%

superforecaster statistician · openai/gpt-5.6-sol

I use the software-ledger base rate of 0.452, with the 31–90-day and threshold classes offering modest upward support. However, the supplied record does not specify the required expansion or show a primary commitment from Artificial Analysis. The rule requires complete published index coverage by a hard date; partial pages miss and ambiguity resolves against the forecast. Publication timing is an additional failure mode. The forecaster’s 0.8–0.9 confidence bin has realized only 0.50 across four cases, so 0.83 appears overconfident absent specific evidence.

57%

geopolitics regulation analyst · openai/gpt-5.6-sol

The 55-day window is plausible for a software-index update, and the threshold and 31–90-day reference classes are relatively favorable. However, the software base rate is only 0.452, while the forecaster’s 0.8–0.9 confidence bin has realized just 0.5. Resolution requires complete published coverage—not an announcement, preview, or partial page—and ambiguity resolves against the forecast. The supplied record also does not specify the exact expansion, increasing verification risk. I therefore discount the stated 0.83 substantially.

Evidence the engine used

direct · undated

AI Model &amp; API Providers Analysis | Artificial Analysis

Cited by semiconductor-supply-chain-analyst: A 55-day window is sufficient for a software-index expansion, and Artificial Analysis can publish without external hardware or OEM dependencies. However, the supplied digest gives no direct evidence of implementation progress, preview results, or a dated