Skip to content

AI Market · Forecast desk · Live · capital

49%

At least two frontier labs publish network isolation or containment requirements for third-party cyber evaluation partners by January 31, 2027.

Probability history

0%25%50%75%100%08-0809-06deadline

Called at 44% on August 8, 2026. Repriced 1 time; now 49%. 147 days to deadline.

How the number was built

Final method: stacked-fixed-weights (resolved n=57 < 80). Each rung is an independent estimate; the final is a stacked combination.

Stated (issue date)44%
Reference class (n=19)65%
Recalibrated53%
Feature model79%
Council34%
Final49%

Drivers

  • Reference class qtype:threshold runs 0.65 (n=19) vs stated 0.44.
  • Hierarchical calibration moves stated 0.44 to 0.53 (lens capital, kappa_stated=24.0, kappa_lens=200.0).
  • Feature model at 0.79: conjuncts=1, named disclosure=False, horizon=176d.
  • Council of 5 at 0.34.

Cautions

  • all 5 votes came from one vendor (openai); extremization disabled, persona diversity only
  • 2 cited URL(s) failed verification and were dropped.

Council

5 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view, aggregated by log odds mean without extremization.

34%

grid and power economist · openai/gpt-5.6-sol

The requirement is unusually specific: public lab policy pages must state network isolation or containment controls applicable to evaluation vendors, and two labs must qualify. Frontier-lab frameworks increasingly mention external evaluations and model security, but detailed vendor controls are usually confined to contracts, security questionnaires, or nonpublic operating procedures. Updating one policy within 147 days is plausible; two qualifying publications require either coordinated standards pressure or a shared security incident. Ambiguity against the forecaster also makes generic sandboxing, model-access, or cybersecurity language insufficient. I am therefore below the forecaster’s 0.44.

34%

semiconductor supply chain analyst · openai/gpt-5.6-sol

The requirement is unusually specific: two separate labs must publicly state network-isolation or containment controls applicable to evaluation vendors. OpenAI, Anthropic, and Google DeepMind periodically revise safety frameworks, creating plausible publication opportunities. However, operational controls for outside evaluators are commonly kept in contracts, security addenda, or confidential onboarding material rather than public policy pages. The short window, two-lab threshold, and ambiguity-against-forecaster rule substantially reduce the chance of a hit. This is below the forecaster’s 0.44.

34%

capital markets credit analyst · openai/gpt-5.6-sol

Frontier labs increasingly publish safety frameworks, but these usually specify outcome-level security principles rather than concrete network-isolation or containment controls imposed on external evaluation vendors. The rule requires qualifying language on official policy pages from two distinct labs; private contract terms, evaluator-authored standards, general sandboxing language, or publication by only one lab fail. Although 147 days permits policy revisions, obtaining two unambiguous public disclosures remains a meaningful conjunctive hurdle, especially under adverse ambiguity resolution. The capital-market series offers little direct signal for this operational-policy event.

31%

superforecaster statistician · openai/gpt-5.6-sol

I use the >90-day base rate of 0.40; the higher capital and threshold rates are less relevant to this narrow policy-disclosure question. Detailed network-isolation terms are often contractual or operational rather than published on lab policy pages. Requiring two separate labs compounds adoption and disclosure-timing risk, and ambiguity about what qualifies as an eval vendor or containment requirement resolves against a hit. A shared safety standard could generate correlated publication, partly offsetting that penalty. The forecaster’s 0.44 confidence band has only one resolved case, so it provides no meaningful calibration adjustment; broader bins show overconfidence rather than underconfidence.

38%

geopolitics regulation analyst · openai/gpt-5.6-sol

The requirement is unusually specific: two distinct frontier labs must publicly state network-isolation or containment controls applicable to evaluation vendors, not merely endorse external evaluations or describe internal model-security practices. Such vendor controls are commonly contractual, audit-based, or confidential, reducing incentives to place operational details on public policy pages. Labs may update safety frameworks around year-end, and one publication is plausible as third-party evaluation expands, but obtaining two qualifying documents within 147 days is materially harder. Ambiguity resolving against the forecaster further lowers the chance that general secure-environment language counts.

Evidence the engine used

direct · undated

Anthropic’s Responsible Scaling Policy \ Anthropic

Cited by grid-and-power-economist: The requirement is unusually specific: public lab policy pages must state network isolation or containment controls applicable to evaluation vendors, and two labs must qualify. Frontier-lab frameworks increasingly mention external evaluations and model security, bu