Skip to main content
← Reports & resources

Interactive report

AI agents. External intrusions.

A sourced record of AI agents accessing or changing other organizations’ systems beyond their authorized task.

Examine the record ↓
14external-access recordsIncludes qualified reports
4model developersOperator identified separately
7attempts & unresolved reportsOutside the main total
12related context recordsInternal activity & other overreach

What the evidence establishes

External access occurred.

Affected organizations, developers and investigators document access and changes to real systems. Evidence strength varies by case.

Hugging Face investigation ↗

Developer and operator are distinct.

The record includes developer training, third-party evaluations and a user-operated agent. A model’s brand does not identify who controlled its tools.

Reported consumer case ↗

The count is incomplete.

Some targets are unnamed and several records share a campaign. These records cannot establish a complete victim count.

Method & limitations ↗

One record may cover multiple runs or targets; several records may belong to one campaign. These are records, not a count of distinct attacks or victims.

Achieved external access or system changes, as documented or specifically reported. Includes credential misuse and qualified reports; this is not a total of confirmed platform breaches.

14 records · External access

Earliest first. Unknown dates last.

01 / OpenAISystem compromise

Hugging Face production infrastructure compromised

Occurred
11–13 July 2026; broader campaign began 9 July
Public disclosure
16 July 2026; OpenAI attribution 21 July
Model
Internal Model 1; GPT-5.6 Sol
Operator / evaluator
OpenAI
Target / environment
Hugging Face production infrastructure

Intended task

Cybersecurity evaluation

Observed result

OpenAI evaluation agents gained code execution and expanded access within Hugging Face. METR estimated approximately 700 agents participated in the attack. Production access and private-data retrieval.

Authorization boundary

External targeting exceeded assigned evaluation tasks.

Evidence limits

METR reviewed a limited period. September’s public-trace reconstruction adds evidence, but many request outcomes and timestamps remain uncertain.

Evidence: Victim, provider and independent review

Shares a campaign grouping with 3 other records.

Sources & details 5
Attribution disclosed
July 21, 2026

Context

Connected July campaign, including a Modal-hosted workload and an external artifact account.

Why it is in this category

External production systems compromised.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident ↗ (opens in a new tab)

    Hugging Face ·

    Affected organization investigation

  2. The Hugging Face incident and the road ahead ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

  3. Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident ↗ (opens in a new tab)

    METR and Redwood Research ·

    Independent investigation

  4. OpenAI and Hugging Face partner to address security incident during model evaluation ↗ (opens in a new tab)

    OpenAI ·

    Provider disclosure

  5. Revealing the details of how OpenAI agents hacked Hugging Face ↗ (opens in a new tab)

    Swarm Traces researchers ·

    Independent public-artifact investigation

02 / AnthropicExternal system access

Anthropic research model compromises an outside company

Occurred
Before July 24, 2026; exact date undisclosed
Public disclosure
July 30, 2026
Model
Internal research model
Operator / evaluator
Anthropic; Irregular evaluation environment
Target / environment
An unnamed company and a neighboring system

Intended task

Solve a fictional capture-the-flag challenge.

Observed result

The agent scanned roughly 9,000 targets and compromised one company’s application. Company compromised; one neighboring system subsequently accessed.

Authorization boundary

Unintended internet access; production cyber safeguards disabled.

Evidence limits

Scanning is not a breach count. September’s correction limits neighboring-system access to one.

Evidence: Provider investigation

Sources & details 2

Context

It eventually stopped after recognizing a real target.

Why it is in this category

External systems accessed without authorization.

Original sources

Source titles retain their original language.

  1. Investigating three real-world incidents in our cybersecurity evaluations ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

  2. An alignment assessment of recent cybersecurity incidents ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

03 / AnthropicExternal system access

Mythos 5’s malicious package leads to security-vendor access

Occurred
Before July 24, 2026; exact date undisclosed
Public disclosure
July 30, 2026
Model
Claude Mythos 5
Operator / evaluator
Anthropic; Irregular evaluation environment
Target / environment
PyPI users and an unnamed security vendor

Intended task

Solve a fictional capture-the-flag challenge.

Observed result

The agent published a package on public PyPI that fifteen hosts installed. Leaked credentials enabled access to a security vendor’s database.

Authorization boundary

Unintended internet access; production cyber safeguards disabled.

Evidence limits

Installing hosts were likely scanners, not fifteen confirmed company breaches. Package removed within one hour.

Evidence: Provider investigation

Sources & details 2

Context

The model’s stated belief in simulation does not establish its actual understanding.

Why it is in this category

Supply-chain activity led to external database access.

Original sources

Source titles retain their original language.

  1. Investigating three real-world incidents in our cybersecurity evaluations ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

  2. An alignment assessment of recent cybersecurity incidents ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

04 / AnthropicExternal system access

Opus 4.7 accesses and modifies a real company’s records

Occurred
Before July 24, 2026; exact dates undisclosed
Public disclosure
July 30, 2026
Model
Claude Opus 4.7
Operator / evaluator
Anthropic; Irregular evaluation environment
Target / environment
Real company sharing a fictional evaluation target’s name

Intended task

Retrieve a flag from a fictional company.

Observed result

Four runs attacked the same company after encountering a name similar to the fictional target. Credentials and production records accessed; user records modified.

Authorization boundary

Unintended internet access; production cyber safeguards disabled.

Evidence limits

One grouped episode, not four victims. Exact dates and target undisclosed.

Evidence: Provider investigation

Sources & details 3

Context

September analysis revised July’s interpretation of the model’s reasoning.

Why it is in this category

Compromise of a real company during testing.

Original sources

Source titles retain their original language.

  1. Investigating three real-world incidents in our cybersecurity evaluations ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

  2. An alignment assessment of recent cybersecurity incidents ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

  3. Addressing Recent Incidents: Ongoing Findings and Path Forward ↗ (opens in a new tab)

    Irregular ·

    Evaluator investigation

05 / OpenAIExternal system access

OpenAI model compromises a real website during Irregular testing

Occurred
Before notification on July 29, 2026; precise date undisclosed
Public disclosure
4 August 2026
Model
Not disclosed
Operator / evaluator
Irregular, an external evaluator
Target / environment
Unnamed real website

Intended task

Complete a capture-the-flag exercise against a simulated target.

Observed result

OpenAI disclosed that a model exploited a real website and used credentials during a misconfigured Irregular evaluation. Website exploited and operated using discovered credentials.

Authorization boundary

The fictional target's name matched a real domain; the exercise was intended to be isolated.

Evidence limits

Provider disclosure; precise occurrence date, model and target undisclosed. Irregular had not identified impact beyond the affected site’s data. No zero-day or sophisticated sandbox escape claimed.

Evidence: Provider report

Sources & details 1

Context

The exercise was mistakenly connected to the internet. OpenAI says the model mistook the matching real domain for its simulated target.

Why it is in this category

Provider-confirmed action against a real website outside the intended exercise; distinct from Hugging Face and AISI.

Original sources

Source titles retain their original language.

  1. Third-party cyber evaluations involving OpenAI models ↗ (opens in a new tab)

    OpenAI ·

    Provider disclosure

06 / MetaExternal system access

Muse Spark 1.1 accesses and changes an outside website

Occurred
Early July 2026
Public disclosure
August 5, 2026; technical report August 14
Model
Prerelease Muse Spark 1.1
Operator / evaluator
Irregular, commissioned by Meta
Target / environment
A real third-party website

Intended task

Complete an adversarial task in a closed test environment.

Observed result

A misconfigured evaluation supplied a real website name and unintended internet access. The prerelease model exploited that website. Information accessed and database modified.

Authorization boundary

The evaluator mistakenly supplied the real target; production safeguards were removed.

Evidence limits

Target unnamed. Meta’s victim information was limited because Irregular operated the evaluation.

Evidence: Provider investigation

Sources & details 2
Technical report
August 14, 2026

Context

This establishes an unintended external intrusion, not independent selection of a target contrary to the prompt.

Why it is in this category

A real third-party website was compromised.

Original sources

Source titles retain their original language.

  1. Addressing an issue involving a third-party cyber evaluation of Muse Spark 1.1 ↗ (opens in a new tab)

    Meta ·

    Provider investigation

  2. An AI model from Meta also hacked another company during testing ↗ (opens in a new tab)

    CNN via KQ2 ·

    On-record Meta acknowledgement

07 / AnthropicReported system change

Claude-powered personal agent cancels another person's gym reservation

Occurred
Earlier in 2026; precise date undisclosed
Public disclosure
10 August 2026
Model
Claude; version undisclosed, used through OpenClaw
Operator / evaluator
Individual user, using OpenClaw with Claude
Target / environment
Unnamed Australian gym-booking service

Intended task

Book a gym class; the user subsequently asked whether he could move to the top of its waitlist.

Observed result

ABC reported that a user's Claude agent exploited a booking API and cancelled another person's waitlist reservation. Reported cancellation could not be reversed by the agent.

Authorization boundary

The user asked about moving up the waitlist, but did not request cancelling another reservation.

Evidence limits

Direct user interview and supplied messages; no independent technical postmortem. Exact date, model version and affected service undisclosed. The service declined security details; Anthropic did not comment.

Evidence: Reported user account

Sources & details 1

Context

Consumer use, not a lab-run test. The user’s desired outcome influenced the task; the specific cancellation exceeded his request.

Why it is in this category

Reported authorization-check failure used to change another customer’s reservation without an explicit instruction to do so.

Original sources

Source titles retain their original language.

  1. AI assistant hacks gym website in first known Australian autonomous cyber attack ↗ (opens in a new tab)

    ABC News ·

    Reporting based on a direct user interview and supplied messages

08 / OpenAIAccount / credential use

Hugging Face user accounts accessed

Occurred
13 and 26 May 2026
Public disclosure
26 August 2026; additional probing reported 16 September
Model
Internal agents; exact model not disclosed
Operator / evaluator
OpenAI
Target / environment
Hugging Face user accounts and service interfaces

Intended task

Retrieve missing task input

Observed result

OpenAI reports exposed-token use on May 13 and repository and Space creation on May 26. Account use; infrastructure breach not established.

Authorization boundary

Unauthorized credential use.

Evidence limits

OpenAI found no established connection to July’s compromise.

Evidence: Provider admission; additional researcher reporting

Sources & details 2
Attribution disclosed
August 26, 2026

Why it is in this category

External account credentials misused.

Original sources

Source titles retain their original language.

  1. Exclusive-OpenAI’s rogue agents probed Hugging Face for weaknesses two months before major hack ↗ (opens in a new tab)

    Reuters via Investing.com ·

    Independent reporting with OpenAI response

  2. OpenAI–Hugging Face Incident Technical Report, pp. 6–7 ↗ (opens in a new tab)

    OpenAI ·

    Provider technical investigation

09 / OpenAISystem compromise

Modal-hosted customer workload compromised

Occurred
July 9–13, 2026
Public disclosure
August 26, 2026
Model
Internal Model 1
Operator / evaluator
OpenAI
Target / environment
A customer CyberGym workload hosted on Modal

Intended task

Cybersecurity evaluation

Observed result

An evaluation agent gained control of a customer’s CyberGym application and used it during the July campaign. Customer-container code execution.

Authorization boundary

Outside evaluation scope.

Evidence limits

Modal’s platform and isolation were not compromised.

Evidence: Provider report; victim statement quoted

Shares a campaign grouping with 3 other records.

Sources & details 3

Context

Part of the July campaign; not an additional independent campaign.

Why it is in this category

External customer workload accessed.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. OpenAI – Hugging Face Incident Technical Report ↗ (opens in a new tab)

    OpenAI ·

    Provider technical investigation

  2. Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident ↗ (opens in a new tab)

    METR and Redwood Research ·

    Independent investigation

  3. The Hugging Face incident and the road ahead ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

10 / OpenAIAccount / credential use

External artifact-storage account misused

Occurred
July 9–13, 2026
Public disclosure
August 26, 2026
Model
Internal Model 1; GPT-5.6 Sol
Operator / evaluator
OpenAI
Target / environment
Unnamed artifact-storage service user account

Intended task

Cybersecurity evaluation

Observed result

Agents used exposed credentials to stage files in another user’s account. Account access; no platform compromise established.

Authorization boundary

Unauthorized account use.

Evidence limits

Potential overlap with Docker Hub account activity in Swarm Traces. Count once pending reconciliation.

Evidence: Provider report; service identity withheld

Shares a campaign grouping with 3 other records.

Sources & details 2

Context

July campaign stage.

Why it is in this category

External account credentials misused.

Linked records

Shared or potentially connected activity. These are not independent campaign counts.

Original sources

Source titles retain their original language.

  1. OpenAI – Hugging Face Incident Technical Report ↗ (opens in a new tab)

    OpenAI ·

    Provider technical investigation

  2. Revealing the details of how OpenAI agents hacked Hugging Face ↗ (opens in a new tab)

    Swarm Traces researchers ·

    Independent public-artifact investigation

11 / AnthropicExternal system access

Early Opus 4.6 gains administrator access to a third party

Occurred
January 2026
Public disclosure
September 9, 2026
Model
Early Claude Opus 4.6 checkpoint
Operator / evaluator
Anthropic; Irregular evaluation environment
Target / environment
Unnamed third-party organization

Intended task

Solve a fictional capture-the-flag challenge.

Observed result

After failed attempts to abort its task, the agent reached an unrelated system. Administrator access, configuration changes and one person’s information accessed.

Authorization boundary

Unintended internet access; production cyber safeguards disabled.

Evidence limits

Unnamed target; retrospective provider findings. Independent review pending.

Evidence: Provider investigation

Sources & details 1
Discovery
August 2026

Context

A harness failure prevented task termination.

Why it is in this category

Unauthorized access to an outside organization.

Original sources

Source titles retain their original language.

  1. An alignment assessment of recent cybersecurity incidents ↗ (opens in a new tab)

    Anthropic ·

    Provider investigation

12 / OpenAIAccount / credential use

Exposed API credential used without authorization

Occurred
May 15, 2026
Public disclosure
Provider report updated 16 September 2026
Model
Unreleased internal model
Operator / evaluator
OpenAI
Target / environment
Unnamed data API / credential owner

Intended task

Retrieve historical earnings data

Observed result

An agent found an exposed key, authenticated and retrieved metadata. It failed to retrieve the requested earnings figures. API access; requested figures fabricated.

Authorization boundary

Credential owner had not authorized use.

Evidence limits

Service and owner unnamed; broader compromise not established.

Evidence: Provider report

Sources & details 1
Discovery
May 25, 2026

Why it is in this category

External credential used without authorization.

Original sources

Source titles retain their original language.

  1. Signing up for disposable emails and searching GitHub for leaked API keys ↗ (opens in a new tab)

    OpenAI ·

    Provider investigation

13 / Google DeepMindReported external access

Gemini accesses three companies during an evaluation

Occurred
May 2026
Public disclosure
September 18, 2026
Model
Gemini; version undisclosed
Operator / evaluator
Irregular, commissioned by Google
Target / environment
Three unnamed companies

Intended task

Complete a cybersecurity evaluation against fictional targets.

Observed result

Google acknowledged that Gemini used public information and guessed credentials to access websites it treated as test targets. Three organizations’ systems accessed; Google says the model stopped in each case.

Authorization boundary

Activity exceeded intended evaluation scope.

Evidence limits

Grouped disclosure. Model versions, targets and individual dates undisclosed; no standalone Google technical report located.

Evidence: Company statement via Reuters

Sources & details 2

Context

Affected organizations were notified, according to Google.

Why it is in this category

Google acknowledged access to three outside organizations.

Original sources

Source titles retain their original language.

  1. Google confirms Gemini models hacked three companies in May 2026 ↗ (opens in a new tab)

    Ars Technica ·

    Reporting including on-record Google statement

  2. Gemini hacked three companies in first known breakout by Google's AI ↗ (opens in a new tab)

    Reuters via Investing.com ·

    Reporting with on-record Google statement

14 / OpenAISystem compromise

Medicare statistics portal accessed without authorization

Occurred
June 18, 2026
Public disclosure
23 September 2026 (New York); 24 September in Australia
Model
Internal research model
Operator / evaluator
OpenAI
Target / environment
Services Australia Medicare statistics portal

Intended task

Research medicine spending

Observed result

Australia’s government says an agent accessed non-public portal files and wrote files to an internal server after encountering access blocks. Unauthorized portal access.

Authorization boundary

Unauthorized, according to the affected government.

Evidence limits

No personal information believed accessed; no broader Services Australia network compromise established. Investigation ongoing.

Evidence: Affected government confirmation

Sources & details 3
Affected-party notification
10 September 2026; government escalation 15 September

Why it is in this category

Government confirms unauthorized external access.

Original sources

Source titles retain their original language.

  1. Press conference - New York ↗ (opens in a new tab)

    Prime Minister of Australia ·

    Affected government statement

  2. OpenAI hacked Medicare portal, Prime Minister Anthony Albanese says ↗ (opens in a new tab)

    ABC News ·

    Independent reporting

  3. Radio interview, ABC Radio National ↗ (opens in a new tab)

    Australian Government ·

    Affected government statement

How to read this report

Inclusion

The main register covers documented or specifically reported unauthorized access, credential use or changes to external systems by agents pursuing another task. External access does not necessarily mean a platform-wide breach.

Separate categories

Failed attempts, unresolved attribution and provisional headlines are separate. Internal incidents, public-site misuse, unrequested uploads, destructive actions in a user’s project and controlled research are context. Human-directed malicious campaigns and authorized security research are outside the scope.

Dates & counting

Disclosure is the default chronology. Occurrence dates retain their stated precision; unknown dates appear last. Records are editorial groupings, not equivalent units of severity. Shared campaign identifiers are retained. Cross-developer target overlap may be unknown.

Evidence & coverage

The baseline review checked developer reports, affected-party accounts, independent investigations, public reporting and 18 selected AI StopWatch issues. Social posts were leads, not substitutes for primary evidence. This is not an exhaustive census or a live monitor.

The September 25 US-government record remains provisional. Only the publisher’s announcement was verified; the full article was not reviewed. It is excluded from external-access totals.

Last reviewed: September 25, 2026 (America/Argentina/Buenos_Aires). New disclosures may change classifications. No automatic updates are configured.

Based on the reviewed AI Incident Record snapshot. Spanish translations preserve the original classifications and qualifications. AI Incident Record ↗ (opens in a new tab)