{
  "question_id": "Q5",
  "slug": "frontier-weight-theft",
  "title": "Theft of closed frontier weights",
  "type": "binary",
  "resolution": {
    "levels": [],
    "values": [],
    "annotations": []
  },
  "forecasts": [
    {
      "question_id": "Q5",
      "forecast_date": "2026-09-14",
      "evidence_cutoff": "2026-09-14",
      "reading_adopted": null,
      "ambiguities": [
        "Permissioned research-weight releases are not explicitly classified as open or closed; I treat intentional external research-weight releases as already released with weights and do not count subsequent redistribution of an authorized external copy as a new unauthorized exit from developer-administered custody.",
        "AI-assisted theft directed by a human qualifies, while autonomous model self-exfiltration without a human theft operation does not.",
        "A sufficiently specific court filing can establish the event without a conviction, but generic allegations of AI intellectual-property theft, endpoint access, or incomplete file exposure are insufficient.",
        "No explicit start date: pre-2026 leaks (Meta LLaMA-1 2023, Mistral miqu 2024) are treated as outside the window and in any case were never confirmed as theft by a covered developer.",
        "Whether a government advisory that alleges theft of 'proprietary AI technology' without naming weights and the model tier would count; I treat it as not qualifying unless weights of a top-two/closed model are specified.",
        "Whether a copy made by an insider to personal storage that is later recovered counts as 'copied out of the developer's control'; I treat it as qualifying per the literal wording.",
        "Whether AI-agent-directed exfiltration without explicit human instruction counts as 'human actor' and what fraction of weight shards counts as 'sufficient to run inference'",
        "A pre-release model that the developer intends to eventually release with open weights (e.g., a next-generation Llama) still counts as 'closed' at the time of theft under clause (i), since it had not yet been released with open weights.",
        "Only successful copying of weight files to non-developer storage counts; thwarted or attempted exfiltrations disclosed by a lab or government do not resolve YES.",
        "A developer-authored threat-intelligence or security report on the developer's own domain (e.g., a Google GTIG report) counts as a 'developer statement' under A2 if it confirms such a theft of that developer's own frontier model.",
        "Whether Meta's 2023 Llama research-preview leak plus takedowns already satisfy the letter of the criteria (adopted: no).",
        "Whether a DMCA takedown notice counts as a court filing (adopted: no).",
        "Whether 'externally deployed' includes research-gated weight distribution rather than product/API deployment.",
        "Whether a government 'theft' accusation that describes API distillation could be misread as weight-file theft (adopted: excluded).",
        "Whether a confirmed infrastructure breach counts without an explicit statement that weight files sufficient to run inference were copied.",
        "Whether any historical leak (e.g., March 2023 LLaMA torrent) already resolves this question; I adopted the reading that it fails both the persistent-qualification test (not among the two most capable externally deployed models at the time) and the exclusion for press-only confirmation, leaving zero prior qualifying events.",
        "Whether reports to the EU AI Office under systemic-risk serious-incident duties count as 'public confirmation'; I treated non-public regulatory filings as not resolving unless publicly released or referenced in a developer/government statement.",
        "Court filings: the criteria list 'court filing' as a resolving source but say 'publicly confirmed'. I assume a complaint/indictment that asserts the weights were copied out qualifies, even if unadjudicated or later dismissed (cf. xAI v. Li; Apple v. OpenAI). A stricter reading requiring an adjudicated finding or an undisputed developer admission would cut my 2031 figure by roughly 5-8 points.",
        "AI-agent theft: in this timeline the most-disclosed intrusions are by the labs' own agents (OpenAI-Hugging Face, July 2026; Anthropic's three eval incidents; UK AISI). The criteria say 'by a human actor', so I exclude fully autonomous model-driven copying, but include human-directed AI-agent operations. If a resolver counted agent self-exfiltration, probabilities would be materially higher (the adjacent Manifold market on model self-exfiltration was at 22% on Aug 12, 2026).",
        "Government source scope: 'government agency' is not limited to the victim's government or to AI safety institutes. I assume a US indictment/advisory/official statement naming a covered developer as victim qualifies (as the Sept 2026 CISA/NSA/FBI advisory did for distillation), and that a Chinese government statement about a Chinese covered developer would too, though the latter is very unlikely to be published.",
        "Whether the A2 dispute rule applies to government/court sources: A2's 14-day non-dispute rule is written for third-party evaluators, AI safety institutes and affected third parties. The question's own resolving-source list includes government indictments/advisories and court filings without that qualifier, so I treat those as qualifying even if the developer disputes them; under the stricter A2 reading a disputed advisory would not resolve.",
        "Accidental exposure (e.g., the March 2026 Anthropic Claude Code packaging leak, or a misconfigured bucket) is excluded: it is neither a human actor copying without authorization nor, in the Claude Code case, weights.",
        "Whether 'closed frontier model' status is judged at theft time for models later open-sourced. I assume yes, per the criteria's '(i) had not been released with open weights at the time of theft'.",
        "Whether existing public material already suffices: xAI's complaint in X.AI Corp. v. Xuechen Li (28 Aug 2025) pleads misappropriation of 'Confidential Information and trade secrets' relating to Grok and notes that trade secrets generally 'protect… model weights, training data', while press retellings (e.g., Feb 2026 Business Insider-derived coverage) say Li took 'confidential Grok materials including weights'. If a resolver reads that as confirmation of weight-file exfiltration, the question is already YES and every horizon is 1.0; press-only sourcing is excluded by the rules, and the operative pleadings describe source code/data, so I do not treat it as resolved — but I load a small probability on it in the first two periods.",
        "Whether a generic government advisory naming Chinese labs for distillation counts (it does not under the explicit 'distillation via API' exclusion), and whether a think-tank/Congressional-witness statement (e.g., 2 Sep 2026 Hudson testimony, AFPI briefs) counts as a 'government agency official statement' (I assume not; only agency documents/indictments/advisories do).",
        "Whether xAI's Grok models satisfy 'closed frontier model' at the mid-2025 time of theft (Grok 3/Grok 4 were closed at that time; Grok 2.5 was open-weighted later), which only matters if the Li allegations are read as covering weights."
      ],
      "key_drivers": [
        "No located public confirmation clearly satisfies all requirements; recent distillation allegations and autonomous evaluation incidents do not qualify.",
        "Rapidly improving cybersecurity capabilities can strengthen human-directed attacks as well as defenses.",
        "One insider's unauthorized inference-sufficient copy can qualify without public release or subsequent deployment.",
        "Repeated generations of closed deployed models and more-capable internal checkpoints create accumulating opportunities across the fixed covered set.",
        "Security hardening, authorized open-weight releases, and alternatives such as distillation reduce theft opportunities and incentives.",
        "Government and court disclosures can reveal older incidents, but secrecy, sealed records, technical ambiguity, and disclosure delays substantially reduce confirmation probability.",
        "Zero publicly confirmed frontier weight thefts in ~4 years despite intense espionage; broad Manifold IP-theft market for OpenAI/Anthropic/DeepMind resolved NO",
        "Developers have strong incentives not to disclose weight theft; government attributions (CISA AA26-251A, Sept 2026) so far target distillation, which is excluded",
        "Rising hazard from US–China AI espionage escalation, insider risk at ten labs, growing attack surface (vendor environments, agentic sandboxes), and new US/Chinese legal tools to prosecute AI-tech theft",
        "Recent incidents (Hugging Face sandbox escapes, Mythos vendor access) were AI-actor or access-level, not human weight exfiltration",
        "Zero qualifying confirmations to date; recent high-profile cases are explicitly excluded (distillation) or non-weights (Ding infrastructure)",
        "Labs at RAND SL2 vulnerable to OC3/OC4 state/insider operations; large insider pool with ~0.5-1% compromise rate",
        "Distillation as cheaper lower-risk substitute reduces incentive for direct weights theft, but rising weights value raises hazard over time",
        "Confirmation lags theft by 1-3 years and developers have incentives to conceal, so forecast is for publication not occurrence",
        "Zero confirmations of closed frontier weight theft in ~4 years of frontier labs being prime nation-state targets; distillation, the only publicly confirmed large-scale theft channel, is explicitly excluded.",
        "Escalating US government willingness to publicize Chinese AI theft (Sept 2026 FBI/NSA/CISA advisory; House Select Committee hearings; threatened sanctions/Entity List actions) raises P(public confirmation | actual theft).",
        "Rapidly improving developer disclosure norms (OpenAI/Anthropic/Meta disclosed sandbox-escape incidents within days in July–Aug 2026; Anthropic publishes threat-intelligence reports) make victim disclosure more plausible than historically.",
        "Metaculus crowd prices ~79% that some frontier lab's weights are stolen or leaked by Jan 2030; this question's stricter confirmation and model-capability bars cut that substantially.",
        "Labs are actively hardening weight security (RAND security-level work, Anthropic's Sept 2026 weights-focused threat model), partially offsetting rising attacker capability and target value.",
        "Confirmation typically lags occurrence by months to years (forensic attribution, declassification, legal process), shifting probability mass to later horizons.",
        "No qualifying confirmation exists today; observed 'theft' is API distillation, which is excluded.",
        "Labs are around RAND SL2-SL3, not SL4, so nation-state theft remains technically plausible, but multi-TB checkpoints plus egress/two-party controls make silent exfil hard.",
        "The resolvable event is official confirmation, which lags and is suppressed by lab, diplomatic, and sources-and-methods incentives.",
        "Highest near-term hazard is insider or leak-then-confirm at weaker-security closed labs (xAI, Meta closed line, ByteDance), plus post-IPO disclosure and DOJ/White House signaling.",
        "Distillation and closing open-weight gap substitute for weight theft; AI-enabled hacking and rising strategic value push the other way.",
        "Zero qualifying public confirmations of closed frontier weight theft through 2023–2026; only near-misses are distillation attacks (excluded) and the Ding supercomputing trade-secret case (not weights).",
        "Intense state espionage on US AI labs (PRC-dominated) with official assessments (IFP, Mandiant, RAND) that weight exfiltration is feasible for well-resourced actors and that no lab meets SL4-5 security standards.",
        "Confirmation pathways are improving: DOJ economic-espionage indictments, multi-agency advisories (Sept 2026 FBI/NSA/CISA distillation advisory as template), Treasury sanctions threats, Anthropic's security-incident disclosure culture, and EU AI Act serious-incident tracking.",
        "Strong confirmation-suppression: labs prefer secrecy, indictments lag events 1–3 years, open-weight alternatives reduce theft incentive, and egress controls on terabyte-scale weight stores are improving.",
        "Hazard set at ~0.8% per quarter in 2026-Q3 rising to ~4% per quarter by 2031, roughly tracking rising model value and regulatory transparency pressure, netting to ~50% by end-2031.",
        "Enormous geopolitical and commercial strategic value of frontier model weights to nation-states (e.g. MSS, SVR) and rogue insiders, creating extreme adversarial pressure.",
        "Broad attack surface covering 10 major developers across the US and China, with dozens of frontier training runs and deployments over 2026–2031.",
        "Stringent resolution gate requiring official public confirmation (developer statement, government indictment/advisory, or court filing), which excludes press leaks, unconfirmed rumors, and API distillation.",
        "Strong commercial and reputational disincentives for developers to voluntarily confirm weight loss, counterbalanced by DOJ criminal indictment patterns for insider economic espionage.",
        "Rapid hardening of frontier lab security perimeters, including ASL-3/ASL-4 egress bandwidth throttling, two-person authorization, and hardware security modules (HSMs).",
        "Zero confirmed cases to date: the International AI Safety Report 2026 states that as of December 2025 there were no confirmed, publicly documented instances of model weight theft, and I found none between January and mid-September 2026.",
        "Disclosure machinery is now unusually strong and keeps strengthening: Anthropic's Aug 2026 Risk Report has a dedicated incident section and explicitly states 'no model weights were accessed' in its April 2026 contractor incident; OpenAI's GPT-6 Astra system card describes purpose-built exfiltration detections; SB 53 (in force Jan 1, 2026) makes unauthorized access/exfiltration of frontier weights a reportable critical safety incident; NY RAISE (72-hour reporting) and IL SB 315 (published audit summaries) start in 2027-28; EU AI Act enforcement began Aug 2026; Amodei's Sept 12, 2026 'We Must Pace the Frontier' commits Anthropic to embedded third-party evaluators whose job includes reporting incidents.",
        "Offense is improving faster than defense: Anthropic's Sept 2026 threat report documents AI-orchestrated kill chains that rewrite their own malware to defeat detections; OpenAI's GPT-6 Astra is the first model at the 'Critical' cyber tier (Sept 3, 2026); RAND says SL5 weight security is 'currently not possible' and Anthropic's ASL-3 explicitly does not cover state-sponsored programs that specifically target it.",
        "Insider/litigation channel is the most likely first confirmation: three frontier-lab trade-secret cases in 18 months (Google's Linwei Ding convicted Jan 29, 2026; xAI v. Xuechen Li; Apple v. OpenAI, July 2026) show that departing-employee IP cases reliably produce public court filings — but every one so far involves code, chip designs or data, not weights.",
        "Government appetite to name and shame is high (CISA/NSA/FBI advisory AA26-251A, Sept 8, 2026; DOJ economic-espionage prosecutions; Senate 'Stealth Stealing' hearing Apr 22, 2026), and frontier weights are now export-controlled items (the June 12, 2026 directive suspending foreign access to Fable 5/Mythos 5), so a theft would be politically useful to publicize.",
        "Offsets that keep the number well below 50%: open-weight models are only ~3 points / ~4 months behind the closed frontier (Epoch AI, mid-2026), which sharply cuts the strategic value of stealing US weights; weights are TB-scale, compartmented, two-party-authorized and egress-monitored; the most likely successful thief (a top-tier state service) is precisely the actor whose success is least likely ever to be publicly confirmed; and the definitional filters (top-2 deployed or more-capable internal model, full inference-sufficient weight set, human actor, not press-only) would knock out most real incidents.",
        "China-based covered developers contribute almost nothing: their frontier models are open-weight (DeepSeek V4, Qwen 3.7, GLM 5.3, Kimi K3) and neither they nor the PRC government would publicize being a victim.",
        "Disclosure, not occurrence, is the binding constraint: no developer, agency or court document yet states that closed frontier weight files left a developer's control (developers explicitly denied it in the OpenAI/Hugging Face and Meta incidents).",
        "The dominant 2026 US government document (FBI/NSA/CISA advisory aa26-251a, 8 Sep 2026) alleges API distillation by DeepSeek/Alibaba/Moonshot/MiniMax/StepFun/Z.AI - a channel this question explicitly excludes - rather than weight exfiltration.",
        "Precedent for insider AI-IP suits is code/designs, not weights (Google/Ding conviction Jan 2026; xAI v. Li pleadings; xAI's OpenAI suits dismissed Feb and Jun 2026), showing both that detection/litigation pathways exist and that companies avoid the 'weights were taken' framing.",
        "Rising state pressure (Treasury sanctions threats, congressional counterintelligence hearings, Chinese state actor abuse of Claude) plus new disclosure duties (EU AI Act enforcement from Aug 2026, California SB 53, industry incident reports, UK AISI incident reporting) raise the rate at which any future theft would surface publicly.",
        "Model weight theft is already occurring commercially outside the frontier-lab reference class (Mandiant, 8 Sep 2026: extortion crews exfiltrating proprietary AI models), evidencing motive and capability, but covered developers retain the strongest security posture.",
        "Loose outside view: aggregated Metaculus/Manifold medians put 'major AI lab announces model weights compromised in a cyberattack' at ~25% by Jan 2027 on wider criteria; adjusted for this question's stricter bar and longer horizon this is consistent with a ~30% five-year cumulative.",
        "Hazard rises slowly with the number of closed frontier models, the number of state and criminal actors targeting them, and tightening mandatory-disclosure regimes."
      ],
      "would_update_on": [
        "A qualifying developer, government, or court disclosure of human-caused exfiltration of complete eligible weights would move relevant horizons toward certainty, subject to the dispute rule.",
        "Documented privileged compromise of a major developer's weight-storage environment with credible copy indicators would raise the end-2027 probability by more than 10 percentage points.",
        "Independently verified, deployed anti-exfiltration protections across several leading closed-model developers could lower the five-year probability by 10–20 percentage points.",
        "A demonstrated shift toward mandatory or routine public reporting of completed weight theft could raise longer-horizon probabilities by more than 10 percentage points.",
        "A binding sustained frontier slowdown with verified hardening, or broad opening of top-model weights, could lower the five-year probability by more than 10 percentage points.",
        "DOJ indictment, CISA/NSA/FBI advisory, or court filing alleging exfiltration of model weights from a named covered developer",
        "Developer security disclosure or undisputed third-party/AISI report confirming weight copying (e.g., further detail on the Claude Mythos vendor incident)",
        "Chinese prosecution under a new AI-technology-theft national-security offence involving a covered Chinese lab's top model",
        "Definitive closure of the 2026 sandbox-escape and Mythos incidents with no exfiltration, and continued exclusive government focus on distillation (would lower the path)",
        "A DOJ indictment, CISA/FBI/NSA advisory, or developer statement alleging actual weight-file exfiltration (not distillation) would raise near-term hazards sharply",
        "Evidence that labs reached RAND SL4/SL5 weights security or that courts ruled distillation-only cases satisfy weights-theft language would move forecast down/up respectively",
        "A credible press report of weights theft (even though non-resolving) followed by government investigation would raise medium-term confirmation probability",
        "A DOJ/FBI indictment or advisory describing actual exfiltration of weight files (not just distillation or credentials) from a covered developer — would move near horizons up by 10+ points.",
        "Any covered developer's incident report or system card confirming weight files were written to non-developer storage — near-certain YES for that quarter.",
        "A court filing (civil suit over stolen weights, or criminal charges against an insider who copied weights).",
        "Evidence that a nation state (e.g., China) is running a copy of a covered developer's closed frontier model, with a US government attribution statement that the developer does not dispute within 14 days.",
        "Stronger disclosure mandates (US AI security legislation, EU AI Act incident-reporting enforcement) — would raise mid/late horizons.",
        "Conversely, credible reporting that labs have achieved RAND SL3+ weight protection and no breach disclosures materialize over 2027 — would push my late horizons down.",
        "Developer, 8-K, or government statement that inference-capable frontier weights were copied off-network.",
        "Unsealed indictment or civil complaint alleging theft of closed frontier weight files.",
        "Verified public appearance of a covered lab's closed frontier checkpoint plus a qualifying confirmation.",
        "Independent assessment that a covered lab has reached RAND SL4 / ASL-4 weight security.",
        "Covered-lab risk report or RSP disclosure of a Critical Safety Incident defined as unauthorized weight exfiltration.",
        "Sustained evidence that open-weight models have matched closed frontier, collapsing the theft prize.",
        "A DOJ indictment, court filing, or official advisory alleging exfiltration of actual model weights rather than distillation or hardware trade secrets.",
        "A covered developer's own security-incident disclosure or threat-intelligence report confirming weight theft.",
        "Credible reports of weight files appearing in the wild (torrents, gray-market offerings), which would force confirmation.",
        "Passage of mandatory public breach-reporting rules for frontier labs (US or EU) that would raise the conversion from occurrence to confirmation.",
        "Decisions by major labs to open more weights, reducing the incentive for theft and hence the hazard.",
        "An unsealed DOJ indictment or FBI complaint charging an actor with unauthorized exfiltration or attempted sale of frontier model weights (+15 to +25 pp).",
        "A public leak of functional frontier weights onto torrent networks or Hugging Face verified by the developer (+20 to +30 pp).",
        "Mandatory US or international statutory reporting regimes requiring public disclosure of critical AI model compromises (+5 to +10 pp).",
        "Widespread shift toward open-source frontier models eliminating the commercial and strategic premium of closed weights (-5 to -10 pp).",
        "Any developer statement, system card, risk report or threat-intelligence report saying that weights (or a checkpoint sufficient to run inference) of a top-tier closed model were accessed, copied or exfiltrated by a person — even a blocked attempt would move me up, and a completed one would resolve it.",
        "A DOJ indictment, CISA/FBI/NSA advisory, ODNI or White House statement, or congressional hearing testimony asserting that a covered developer's frontier weights were taken; the first California OES aggregated critical-safety-incident report (statutorily due from Jan 1, 2027) and the first published Illinois SB 315 audit summaries (from Jan 1, 2028) are scheduled windows I will check.",
        "A court filing by a covered developer alleging an employee/contractor copied frontier model weights (the xAI v. Li and Apple v. OpenAI pattern), or an SEC Item 1.05 8-K from Alphabet, Meta or Microsoft describing loss of frontier model IP.",
        "Evidence of a step-change in weight security (labs reaching independently verified RAND SL3/SL4, checkpoint encryption plus air-gapping becoming standard, staff siloing per Amodei's pacing plan) — would move me down 5-10 points across horizons; conversely disclosure of a serious near-miss at weight storage, or of insiders with standing access to multiple top models, would move me up.",
        "A credible dark-web/forum claim of possessing frontier weights followed by a developer response (confirmations resolve YES; denials or silence plus press-only reporting do not), and any METR/UK AISI/US CAISI report describing weight exfiltration at a lab that the lab does not dispute within 14 days.",
        "Geopolitical escalation that makes weight theft a top-priority state operation (kinetic US-China conflict, Entity List designations of DeepSeek et al. actually being published, or a US decision to place frontier weights under formal national-security custody), which would raise both the attempt rate and the disclosure rate.",
        "A DOJ indictment, amended complaint, or developer statement that specifically names weight files (e.g., a Xuechen Li indictment or xAI filing saying Grok weights were uploaded to personal storage) - would lift the nearest horizon to 0.5-0.7 and raise the whole curve.",
        "A CISA/NSA/UK AISI advisory, EU AI Office finding, or NIST/CAISI report describing a successful model-weight exfiltration (rather than targeting or distillation) at a covered developer.",
        "Evidence that extortion or espionage crews posted verifiable weights of a covered developer's closed model, forcing a developer confirmation or an undisputed third-party attribution.",
        "Explicit developer statements that recent investigated incidents did NOT involve loss of weights (raising P(occurrence), or conversely, evidence of routine quiet settlements suggesting confirmations are systematically suppressed - which would cut the five-year figure below 0.20).",
        "Clarification from the resolution process that 'trade secrets including weights' allegations suffice, which would immediately set p_yes to ~1.0 for all horizons."
      ],
      "forecasts": [
        {
          "period_end": "2026-09-30",
          "p_yes": 0.0082
        },
        {
          "period_end": "2026-12-31",
          "p_yes": 0.0286
        },
        {
          "period_end": "2027-03-31",
          "p_yes": 0.0526
        },
        {
          "period_end": "2027-06-30",
          "p_yes": 0.0769
        },
        {
          "period_end": "2027-09-30",
          "p_yes": 0.1002
        },
        {
          "period_end": "2027-12-31",
          "p_yes": 0.125
        },
        {
          "period_end": "2028-03-31",
          "p_yes": 0.1512
        },
        {
          "period_end": "2028-06-30",
          "p_yes": 0.1765
        },
        {
          "period_end": "2028-09-30",
          "p_yes": 0.2022
        },
        {
          "period_end": "2028-12-31",
          "p_yes": 0.228
        },
        {
          "period_end": "2029-03-31",
          "p_yes": 0.2536
        },
        {
          "period_end": "2029-06-30",
          "p_yes": 0.2803
        },
        {
          "period_end": "2029-09-30",
          "p_yes": 0.3053
        },
        {
          "period_end": "2029-12-31",
          "p_yes": 0.3295
        },
        {
          "period_end": "2030-03-31",
          "p_yes": 0.3549
        },
        {
          "period_end": "2030-06-30",
          "p_yes": 0.3789
        },
        {
          "period_end": "2030-09-30",
          "p_yes": 0.4023
        },
        {
          "period_end": "2030-12-31",
          "p_yes": 0.4243
        },
        {
          "period_end": "2031-03-31",
          "p_yes": 0.4469
        },
        {
          "period_end": "2031-06-30",
          "p_yes": 0.4678
        },
        {
          "period_end": "2031-09-30",
          "p_yes": 0.4878
        }
      ]
    }
  ]
}