Skip to content

AI Market · Forecast desk · Live · hardware

44%

SemiAnalysis publishes AgentX v3 multi-turn benchmark results for OpenAI Jalapeño on production-representative agentic traces, with methodology comparable to Vera Rubin NVL72 AgentX runs cited by NVIDIA, by October 31, 2026.

Probability history

0%25%50%75%100%08-2909-06deadline

Called at 36% on August 29, 2026. Repriced 1 time; now 44%. 55 days to deadline.

How the number was built

Final method: stacked-fixed-weights (resolved n=57 < 80). Each rung is an independent estimate; the final is a stacked combination.

Stated (issue date)36%
Reference class (n=8)39%
Recalibrated52%
Feature model63%
Council32%
Final44%

Drivers

  • Reference class qtype:timing runs 0.39 (n=8) vs stated 0.36.
  • Hierarchical calibration moves stated 0.36 to 0.52 (lens hardware, kappa_stated=24.0, kappa_lens=200.0).
  • Feature model at 0.63: conjuncts=1, named disclosure=False, horizon=63d.
  • Council of 5 at 0.32.

Cautions

  • all 5 votes came from one vendor (openai); extremization disabled, persona diversity only

Council

5 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view, aggregated by log odds mean without extremization.

33%

grid and power economist · openai/gpt-5.6-sol

This requires an exact publication outcome: Jalapeño must become benchmarkable and SemiAnalysis must publish it specifically in AgentX v3 within 55 days. A model announcement, API release, preview, or result on another benchmark would all miss. The timing reference class is only 0.389, and the additional dependency on SemiAnalysis testing and publication pushes the estimate lower. Hardware and 31–90-day base rates are less diagnostic because they include less document-specific outcomes. The supplied public-series data provides no direct evidence of scheduled Jalapeño availability or an AgentX v3 update.

31%

semiconductor supply chain analyst · openai/gpt-5.6-sol

This requires a specific publisher, benchmark version, model name, and publication by a near-term deadline. Jalapeño must become testable early enough for SemiAnalysis to complete multi-turn runs and publish AgentX v3 results; an OpenAI announcement or results on another benchmark would not qualify. The 31–90-day and hardware base rates are favorable, but the more relevant timing base rate is only 0.389, and ambiguity resolves against a hit. With no pertinent public-series evidence of imminent benchmark availability, I place this modestly below the forecaster’s 0.36.

31%

capital markets credit analyst · openai/gpt-5.6-sol

The strict trigger requires SemiAnalysis to publish a specific AgentX v3 result for a specifically named OpenAI model within 55 days. This depends on Jalapeño being available to the benchmark team, AgentX v3 testing being completed, and publication occurring by the deadline. Benchmark-version and model-timing slippage are common, while results on any other benchmark receive no credit. The timing reference class is weak at 0.389, and ambiguity resolves against a hit. Hardware and 31–90-day base rates are less diagnostic than the exact publication requirement.

31%

superforecaster statistician · openai/gpt-5.6-sol

I use the timing-question base rate of 0.389 (n=8), rather than the broader hardware or 31–90-day rates. I adjust downward because resolution requires one exact benchmark/version and named system to appear in SemiAnalysis’s publication by a hard date; related benchmarks, testing progress, or later publication do not count. Access, successful multi-turn testing, editorial scheduling, and explicit inclusion each add failure risk despite the nominally single-leg classification. The digest provides no specific evidence of a scheduled Jalapeño result. The forecaster’s sub-0.50 calibration sample is only one case, so it supplies no reliable correction.

34%

geopolitics regulation analyst · openai/gpt-5.6-sol

The 55-day window is plausible for a benchmark update, and the ledger’s 31–90-day and hardware base rates are favorable. However, resolution requires the exact combination of AgentX v3, multi-turn results, and Jalapeño—not merely a model launch or another benchmark. This effectively depends on timely model access, completed evaluation, and SemiAnalysis publication before a hard deadline. The timing reference class is only 0.389, and strict ambiguity treatment increases the chance that near-matching coverage fails. With no concrete publication schedule in the supplied evidence, I place it modestly below that timing base rate.

Evidence the engine used