Skip to content

Radar · AGI and capabilities · T2 · 2029 · CALL

AI solves ten FrontierMath open problems by 2028

Epoch AI's FrontierMath Open Problems page credits AI systems with verified solutions to at least 10 of the listed unsolved research-level problems on or before 2028-12-31.

CALLupindicators pendingregistered 2026-09-08Epoch AI

ClaimEpoch AI's FrontierMath Open Problems page credits AI systems with verified solutions to at least 10 of the listed unsolved research-level problems on or before 2028-12-31.
Consensus (implied)35%implied from Epoch AI FrontierMath Open Problems page and Artificial Analysis GPQA Diamond results · 2026-09-07
Distance+1.24log-odds · clearly above consensus
My confidence65%80% CI 5080%
Engine42%-23 pts vs me · council-only:log-odds-mean
Falsifies ifOn 2027-12-31 Epoch credits AI with fewer than 4 verified open-problem solutions.
HorizonDecember 31, 2028846 days · by end-2029 · milestone ladder

Why it matters

Multiple-choice science benchmarks are exhausted; GPQA Diamond fell before this book was written. Solving research problems nobody has solved is the first rung that cannot be reached by memorization. Ten verified solutions would mean models are producing new mathematics at a rate that changes how research labs staff.

Probability over time

0%25%50%75%100%09-0709-0709-08deadline

Registered at 65% on September 8, 2026. Engine repriced 2 times; now 42%.

Milestone ladder

Dated rungs. Each is scored on its own; the thesis does not get credit for the ladder until the rungs land.

0%50%100%2027-06-30m175%2027-12-31m270%2028-12-31m365%

filled bar · my probabilityhollow dot · engineamber date · due, awaiting adjudication

m1 · 2027-06-30 · 75% · Epoch credits AI with at least 4 verified open-problem solutions.

m2 · 2027-12-31 · 70% · Epoch credits AI with at least 6 verified open-problem solutions.

m3 · 2028-12-31 · 65% · Epoch credits AI with at least 10 verified open-problem solutions.

Leading indicators

Registered thresholds. Status is computed from the latest public reading.

FrontierMath open problems marked solved by AIno data

count

42no data yet

on track >= 4 · off <= 2epoch

connector returned no usable reading · checked 2026-09-07

What would move me

Would raise my number

  • A frontier lab publishes a verified open-problem solution as a headline result in 2027.
  • Epoch verifies a third and fourth solution before mid-2027.
  • Mathematicians submit new open problems faster than Epoch can verify solutions, signalling demand.

Would cut it

  • Epoch tightens verification and reverses a credited solution.
  • Author verification of submitted solutions takes longer than six months.
  • Labs stop evaluating on FrontierMath after saturating Tiers 1-3.

Council · 42%

6 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view.

43%

grid-and-power-planner · openai/gpt-5.6-sol

Starting from 2, this requires eight additional verified solves in 28 months, about 3.4 per year. My reference class is early research-benchmark progress: capability gains are lumpy, while verification and publication lag model performance. A Poisson-style extrapolation from roughly two solves during 2026 lands near, but below, the threshold. Tier-4 saturation is supportive, yet transfer from bounded benchmark questions to genuinely open problems is weak, and the remaining set may be selected for resistance. The literal rule also requires Epoch—not merely papers or announcements—to mark ten qualifying, non-human-assisted solutions; page retirement adds another failure mode. Thus 0.65 appears optimistic.

41%

semiconductor-supply-chain-analyst · openai/gpt-5.6-sol

The reference class is sparse early successes on newly launched, adversarial research benchmarks, where closed-benchmark saturation transfers poorly to genuinely open problems. From 2 solved, the thesis needs 8 additional verified solves in roughly 28 months—about 3.4 annually—with harder-tail depletion and review lag. Capability acceleration and a pool of roughly 48 remaining problems make that plausible, but the strict rule excludes partial or human-assisted work. Resolution also conjunctively requires Epoch’s designation, qualifying solution dates, and continued page tracking. I estimate about 0.44 conditional on continued tracking, reduced to 0.41 overall, below the stated 0.65.

38%

enterprise-cio · openai/gpt-5.6-sol

Starting from 2, the thesis requires 8 additional verified solves in roughly 28 months, about 3.4 per year. My reference class is transitions from benchmark saturation to independently verified research results: benchmark scores can improve quickly, but novel proofs face search, evaluation, attribution, and publication lags. A count model centered on 2–3 accepted solves in 2027 and 3–4 in 2028 puts the threshold in the upper tail even with acceleration. Strict exclusion of human-assisted and under-review work, plus page-continuity risk, lowers the odds further. Tier 4 saturation is supportive but not a reliable linear predictor of autonomous open-problem closure.

44%

credit-analyst · openai/gpt-5.6-sol

Using sparse, independently verified research-mathematics breakthroughs as the reference class, the target requires eight additional credited solutions in about 2.3 years, or roughly 3.5 annually. A Poisson-style baseline near three annual additions puts the threshold below even odds. Rapid capability gains create substantial upside, but benchmark saturation does not reliably translate into autonomous theorem solving. Correlated problem difficulty, verification delays, exclusion of human-assisted work, and the page-continuity condition reduce resolution probability. With no indicator update beyond the baseline count of two, I place the claim below the forecaster’s 0.65.

47%

geopolitics-policy-analyst · openai/gpt-5.6-sol

Starting from 2, the thesis needs 8 additional verified solutions in about 28 months. A Poisson-style reference class for rare, independently verified research breakthroughs, calibrated loosely to two arrivals during 2026, puts the required pace near the upper-middle of the plausible range. Rapid benchmark gains support an accelerating arrival rate, but Tier 4 saturation is weaker evidence for solving genuinely open problems. Epoch’s verification delay, exclusion of human-assisted or under-review work, and the requirement that the page remain active all reduce resolution probability. The missing indicator update adds uncertainty rather than positive evidence. This makes 0.65 too optimistic under strict resolution.

38%

superforecaster-statistician · openai/gpt-5.6-sol

Base rate: I use a 0.30 chance that a curated set of research-level open problems reaches 20% verified autonomous completion within roughly 28 months after only two initial machine solves. The target needs eight additional successes among about 48, while review latency, exclusion of human-assisted work, and page-retirement risk all reduce resolution odds. Rapid FrontierMath Tier 4 saturation is positive, but benchmark performance is a weaker reference class than novel, author-verified research, and the connector supplies no confirming trajectory. Capability acceleration raises my estimate above the base rate, but 0.65 underweights threshold and verification risk.