Top SWE-bench Pro public resolved rateno data
— percent
on track >= 86 · off < 83manual ↗
connector returned no usable reading · checked 2026-09-07
Radar · AGI and capabilities · T1 · 2027 · CALL
A model posts a resolved rate of at least 90.0% on the SWE-bench Pro public leaderboard (Scale AI, public problem set) on or before 2027-12-31.
This is the first rung of the capability ladder: can a model close real, uncontaminated repository issues at the rate a strong senior engineer would. Passing it moves the enterprise question from whether coding agents work to how much software headcount they displace, which is what rung four measures.
Registered at 80% on September 8, 2026. Engine repriced 2 times; now 68%.
Registered thresholds. Status is computed from the latest public reading.
Top SWE-bench Pro public resolved rateno data
— percent
on track >= 86 · off < 83manual ↗
connector returned no usable reading · checked 2026-09-07
Artificial Analysis Terminal-Bench 2.1 top scoreno data
— percent
on track >= 85 · off < 78artificial-analysis ↗
connector returned no usable reading · checked 2026-09-07
Would raise my number
Would cut it
6 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view.
74%
grid-and-power-planner · openai/gpt-5.6-sol
Using the reference class of public coding benchmarks already above 80%, I estimate a 0.79 chance of gaining the remaining 9.7 points within 16 months. The supplied rise from roughly 23% to 80.3% supports rapid improvement, and accepted scaffolds plus a public problem set increase benchmark-specific optimization. However, progress usually slows near saturation because remaining tasks are unusually difficult or evaluation-sensitive. I apply about a 0.94 probability that a qualifying result is actually dated, listed under Scale’s standard evaluation, and the leaderboard or same-metric successor remains available. Multiplying gives approximately 0.74, below the forecaster’s 0.80.
72%
semiconductor-supply-chain-analyst · openai/gpt-5.6-sol
My reference class is threshold crossings on rapidly improving but saturating software-agent benchmarks, anchored to the 0.65 structural prior. Moving from 80.3% to 90.0% requires roughly halving the remaining error rate over 16 months. New model generations, inference scaling, and accepted scaffolds make that plausible, while the public set permits optimization and possible contamination. However, the hardest residual tasks can create a long tail, and the observed one-year gain should not be extrapolated linearly. I estimate about 0.77 for a qualifying technical result, discounted for leaderboard retirement, methodology changes, submission delays, or failure to list the result by the deadline.
68%
enterprise-cio · openai/gpt-5.6-sol
My reference class is frontier software-engineering benchmarks after broad lab adoption: headline scores often saturate within 12–24 months, especially when public tasks and accepted scaffolds permit targeted optimization. Moving from 80.3 to 90.0 requires eliminating roughly half the remaining failures, harder than a straight-line extrapolation suggests because of irreducible, flaky, or environment-sensitive tasks. The 16-month window still allows multiple model and agent generations. I discount for diminishing returns, absent confirming indicators, leaderboard submission discretion, and the conjunctive requirement that Scale continue listing the qualifying result. This is above the 0.65 structural prior but below the stated 0.80.
68%
credit-analyst · openai/gpt-5.6-sol
My reference class is public coding benchmarks after frontier-lab adoption: rapid early gains followed by sharply diminishing returns near saturation. From 80.3%, reaching 90% requires halving the unresolved-task rate, materially harder than adding 9.7 linear points. Still, 16 months, public-task optimization, and accepted tool scaffolds make crossing more likely than not. I estimate a 0.62 chance of crossing in the next 12 months and a 0.24 conditional chance in the final four, then discount roughly 0.03 for leaderboard retirement, methodology, or qualifying-entry risks. The 0.65 structural prior is consistent; without fresh indicators, 0.80 is too confident.
72%
geopolitics-policy-analyst · openai/gpt-5.6-sol
Reference class: frontier software-engineering benchmarks after broad lab adoption, which typically show rapid early gains followed by diminishing returns near saturation. Advancing from 80.3 to 90.0 within roughly 16 months is plausible because accepted tool scaffolds, public tasks, and competition among several labs create repeated attempts. The prior year’s unusually large gain warrants an uplift from the 0.65 structural prior. I discount for increasingly difficult residual tasks, evaluation variance, possible rule tightening, and literal-resolution risk from leaderboard retirement or metric replacement. A roughly 0.62 annual success hazard compounded across 1.32 years yields about 0.72, below the forecaster’s 0.80.
70%
superforecaster-statistician · openai/gpt-5.6-sol
Base rate: 0.65 from the supplied reference class for threshold questions. I adjust upward because the stated leaderboard gain from roughly 23% to 80.3% in one year leaves only 9.7 points over nearly 16 months, and accepted scaffolds broaden the paths to success. I limit that adjustment because the remaining improvement requires about halving the current error rate, benchmark progress commonly slows near saturation, and both qualifying performance and timely listing on the named public leaderboard are required. Retirement without a same-metric successor also causes failure. With no fresh indicator data, the forecaster’s 0.80 relies too heavily on straight-line extrapolation.
65% from reference-class:qtype:threshold. ledger base rate, n=19, horizon 480d; the ledger has no multi-year history