Skip to content

Radar · AGI and capabilities · T2 · 2029 · WARNING

METR does not certify a one-month agent horizon

METR does not publish a 50%-success time-horizon point estimate of 160 work-hours (one working month) or more for any publicly released model on its Time Horizons page on or before 2028-12-31.

WARNINGdownindicators pendingregistered 2026-09-08METR

ClaimMETR does not publish a 50%-success time-horizon point estimate of 160 work-hours (one working month) or more for any publicly released model on its Time Horizons page on or before 2028-12-31.
Consensus (implied)35%implied from METR, Time Horizons page and methodology note (2026-03-20); Manifold, "Best METR 50% Time Horizon in 2026" · 2026-09-07
Distance+1.24log-odds · clearly above consensus
My confidence65%80% CI 5080%
Engine56%-9 pts vs me · council-only:log-odds-mean
Falsifies ifMETR publishes a 50% time horizon of 160 hours or more for a public model on or before 2028-12-31, or publishes a task suite with multi-week human baselines before 2027-12-31 (which would make certification in 2028 likely).
HorizonDecember 31, 2028846 days · by end-2029 · milestone ladder

Why it matters

A month-long autonomous horizon is the rung where agents stop being tools and become staff. Everything downstream, including the labor-market rung and the capacity to run always-on agents per employee, depends on whether this is reached in 2028 or measured only in retrospect. The thesis bets that measurement lags capability, which matters for how anyone reads AGI timelines.

Probability over time

0%25%50%75%100%09-0709-0709-08deadline

Registered at 65% on September 8, 2026. Engine repriced 2 times; now 56%.

Milestone ladder

Dated rungs. Each is scored on its own; the thesis does not get credit for the ladder until the rungs land.

0%50%100%2027-06-30m160%2027-12-31m250%2028-12-31m335%

filled bar · my probabilityhollow dot · engineamber date · due, awaiting adjudication

m1 · 2027-06-30 · 60% · METR publishes a 50% time horizon of at least 40 hours for a public model.

m2 · 2027-12-31 · 50% · METR publishes a task suite with human-baseline tasks longer than 80 hours.

m3 · 2028-12-31 · 35% · METR publishes a 50% time horizon of at least 160 hours for a public model (rung 3 achieved).

Leading indicators

Registered thresholds. Status is computed from the latest public reading.

Longest METR 50% time horizon for a public modelno data

hours

4060no data yet

on track < 40 · off >= 60manual

connector returned no usable reading · checked 2026-09-07

Longest human-baseline task in METR's suiteno data

hours

80200no data yet

on track < 80 · off >= 200manual

connector returned no usable reading · checked 2026-09-07

What would move me

Would raise my number

  • METR revises its doubling-time estimate longer than 8 months.
  • METR publishes that 50% horizons plateau near 30-50 hours across two model generations.
  • Frontier labs stop granting METR pre-release access, delaying public estimates.

Would cut it

  • METR releases a task suite with multi-week human baselines in 2027.
  • A public model reaches 40 hours before mid-2027.
  • Independent replications (Epoch, Apollo) publish horizon estimates consistent with METR's.

Council · 56%

6 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view.

47%

grid-and-power-planner · openai/gpt-5.6-sol

I use rapidly iterated AI benchmarks facing measurement saturation as the reference class. Moving from 17.4 to 160 hours requires about 3.2 doublings; the quoted 4–7-month trend makes underlying capability by end-2028 plausible. Certification is harder: METR must add sufficiently long tasks, obtain costly human baselines, evaluate a public model, deem the result reliable and in-range, and publish the point estimate. A longer task suite alone does not resolve MISS. With no new indicator data, I assign certification hazards of roughly 0.03 in late 2026, 0.20 in 2027, and 0.38 in 2028. Compounding, plus about 0.02 risk that publication ceases, leaves 0.47 for HIT.

54%

semiconductor-supply-chain-analyst · openai/gpt-5.6-sol

Starting from the supplied 17.4-hour estimate, 160 hours requires about 3.2 doublings. At METR’s cited 4–7-month pace, raw extrapolation crosses the threshold around late 2027 to mid-2028. Capability alone is insufficient, however: METR must extend its roughly 16-hour measurement ceiling, validate long-baseline tasks, test a public model, and publish an in-range 50% point estimate by the cutoff. Using METR benchmark-refresh cycles and capability extrapolations as the reference class, I assign conditional certification hazards near 0.18 through 2027 and 0.31 in 2028, compounded, plus a small adverse publication-discontinuation risk. Non-certification remains slightly more likely.

62%

enterprise-cio · openai/gpt-5.6-sol

I use the reference class of frontier-benchmark transitions: capability often outruns a benchmark, while redesign, validation, and publication take roughly 1–2 years. A 17.4-to-160-hour increase is about 3.2 doublings, so the quoted 4–7-month trend makes underlying capability plausible before 2029. But certification is conjunctive: METR must deploy longer-baseline tasks, obtain a reliable point estimate, test a publicly released model, and publish by the deadline. I assign about 0.35 cumulative probability to that sequence, plus roughly 0.03 to METR ceasing publication, yielding 0.62 for HIT. This is slightly below the forecaster’s 0.65 because over two budget years a suite expansion is quite feasible.

58%

credit-analyst · openai/gpt-5.6-sol

I use METR’s own frontier-model horizon progression and benchmark-refresh lag as the reference class. The supplied 4–7-month doubling trend can reach 160 hours before the cutoff, but certification additionally requires a substantially extended task suite, reliable in-range estimation, a public model, and timely publication. Treating qualifying certification as a rising hazard averaging roughly 0.20 annually over the remaining 2.3 years gives about 0.40 cumulative risk; a small probability that METR stops publishing also resolves against the statement. Missing indicator data provides no basis to raise confidence. The conjunction therefore makes HIT modestly more likely than MISS, but below the forecaster’s 0.65.

56%

geopolitics-policy-analyst · openai/gpt-5.6-sol

Using METR’s reported 4–7-month doubling trend as the capability reference class, moving from 17.4 to 160 hours requires about 3.2 doublings, implying a raw crossing around late 2027 to mid-2028. That weighs against the thesis. But resolution requires a conjunction: a public model, an upgraded suite capable of reliable month-scale measurement, a qualifying 50% point estimate, and publication by the cutoff. Benchmark redesign and long human baselines can lag capability by years; the current >16-hour unreliability warning is therefore material. With 27 months remaining, I put certification slightly below even odds. The indicators provide no update.

57%

superforecaster-statistician · openai/gpt-5.6-sol

Base rate: I use 0.45 for a fast-moving capability benchmark to publish a ninefold threshold crossing within 27 months when its fitted trend projects crossing inside that window (about 0.26 annual hazard). The 17.4-hour baseline and 4–7-month doubling trend increase certification risk. However, a MISS requires several timed legs: a substantially expanded task suite, a qualifying public model, a reliable 50%-success estimate, and METR publication by the deadline. The current ~16-hour reliability ceiling makes that measurement chain a major constraint. METR discontinuation adds a small adverse tail. With no new indicator data, 0.57 is below the forecaster’s 0.65.