Skip to content

AI Market · Forecast desk · Live · software

40%

A frontier lab publishes turn count or task-completion cost as a headline metric alongside benchmark scores in an official model card or launch post by January 31, 2027.

Probability history

0%25%50%75%100%08-1509-06deadline

Called at 26% on August 15, 2026. Repriced 1 time; now 40%. 147 days to deadline.

How the number was built

Final method: stacked-fixed-weights (resolved n=57 < 80). Each rung is an independent estimate; the final is a stacked combination.

Stated (issue date)26%
Reference class (n=8)39%
Recalibrated50%
Feature model56%
Council30%
Final40%

Drivers

  • Reference class qtype:timing runs 0.39 (n=8) vs stated 0.26.
  • Hierarchical calibration moves stated 0.26 to 0.50 (lens software, kappa_stated=24.0, kappa_lens=200.0).
  • Feature model at 0.56: conjuncts=1, named disclosure=False, horizon=169d.
  • Council of 5 at 0.30.

Cautions

  • all 5 votes came from one vendor (openai); extremization disabled, persona diversity only

Council

5 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view, aggregated by log odds mean without extremization.

29%

grid and power economist · openai/gpt-5.6-sol

Multiple frontier-lab model releases are likely within 147 days given the digest’s frequent 2026 release cadence. However, model cards usually headline capability, safety, latency, or token-efficiency results—not turn count or end-to-end task-completion cost. Agentic evaluations create a plausible path because longer trajectories make cost and turns decision-relevant, but merely including either metric in a table, appendix, or footnote will not resolve positively. The undefined boundary of “frontier lab” and the ambiguity-against-forecaster rule further reduce the chance. I am slightly above the prior 0.26 because numerous likely release opportunities provide multiple shots at adoption.

29%

semiconductor supply chain analyst · openai/gpt-5.6-sol

Agentic evaluations increasingly expose inefficiency: two models can have similar completion rates but materially different tool-call counts, turns, and inference costs. Multiple active frontier labs and likely year-end model releases create several opportunities. However, model cards still foreground capability and safety benchmarks; pricing, token use, and task cost usually appear in API documentation, benchmark methodology, or footnotes. “Task-completion cost” also requires standardized task and pricing assumptions, while turn count can be harness-dependent. The strict requirement that either metric be a headline model-card field, plus ambiguity resolving against the forecast, keeps the probability below the software and timing base rates.

31%

capital markets credit analyst · openai/gpt-5.6-sol

Several frontier labs are likely to release models before the deadline, and agentic evaluations increasingly track interaction length and end-to-end economics. However, model cards still emphasize capability, safety, latency, and benchmark accuracy; turn count is usually supporting evaluation detail, while task-completion cost is complicated by pricing, scaffolding, and token-budget assumptions. The rule further requires the metric to be a headline field in the model card itself, excluding blog emphasis or footnotes. These constraints keep the probability below the software and timing base rates, though modestly above the forecaster’s 0.26.

29%

superforecaster statistician · openai/gpt-5.6-sol

I use the 0.40 base rate for >90-day questions, consistent with the 0.389 timing rate. Multiple frontier labs and two qualifying metric types provide several paths to a hit, while the digest indicates continued model-release activity. However, the rule requires the metric in a model card and as a headline field by a hard date; ordinary benchmark tables, launch posts, or footnotes fail. “Frontier lab” and “headline” also create resolution risk under the adverse ambiguity policy. The forecaster’s low-confidence calibration has only one observation, while broader results show overconfidence, so it offers no basis for raising this much.

30%

geopolitics regulation analyst · openai/gpt-5.6-sol

Multiple frontier labs are likely to release models within 147 days, and agentic evaluation increasingly makes turns and end-to-end task cost decision-relevant. However, model cards still foreground capability, safety, context length, and benchmark accuracy; turns usually appear in evaluation methodology, while pricing is published separately. The strict requirement that either measure be a headline field—not a table detail or footnote—substantially reduces the chance of a hit. The release-cadence digest raises opportunity count but provides no evidence that headline model-card conventions have shifted.

Evidence the engine used

direct · undated

Models – Hugging Face

Cited by grid-and-power-economist: Multiple frontier-lab model releases are likely within 147 days given the digest’s frequent 2026 release cadence. However, model cards usually headline capability, safety, latency, or token-efficiency results—not turn count or end-to-end task-completion cost. Agent