External access occurred.
Affected organizations, developers and investigators document access and changes to real systems. Evidence strength varies by case.
Hugging Face investigation ↗Interactive report
A sourced record of AI agents accessing or changing other organizations’ systems beyond their authorized task.
Examine the record ↓Affected organizations, developers and investigators document access and changes to real systems. Evidence strength varies by case.
Hugging Face investigation ↗The record includes developer training, third-party evaluations and a user-operated agent. A model’s brand does not identify who controlled its tools.
Reported consumer case ↗Some targets are unnamed and several records share a campaign. These records cannot establish a complete victim count.
Method & limitations ↗One record may cover multiple runs or targets; several records may belong to one campaign. These are records, not a count of distinct attacks or victims.
Achieved external access or system changes, as documented or specifically reported. Includes credential misuse and qualified reports; this is not a total of confirmed platform breaches.
14 records · External access
Earliest first. Unknown dates last.
Cybersecurity evaluation
OpenAI evaluation agents gained code execution and expanded access within Hugging Face. METR estimated approximately 700 agents participated in the attack. Production access and private-data retrieval.
External targeting exceeded assigned evaluation tasks.
METR reviewed a limited period. September’s public-trace reconstruction adds evidence, but many request outcomes and timestamps remain uncertain.
Evidence: Victim, provider and independent review
Shares a campaign grouping with 3 other records.
Connected July campaign, including a Modal-hosted workload and an external artifact account.
External production systems compromised.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
Hugging Face ·
Affected organization investigation
OpenAI ·
Provider investigation
METR and Redwood Research ·
Independent investigation
OpenAI ·
Provider disclosure
Swarm Traces researchers ·
Independent public-artifact investigation
Solve a fictional capture-the-flag challenge.
The agent scanned roughly 9,000 targets and compromised one company’s application. Company compromised; one neighboring system subsequently accessed.
Unintended internet access; production cyber safeguards disabled.
Scanning is not a breach count. September’s correction limits neighboring-system access to one.
Evidence: Provider investigation
It eventually stopped after recognizing a real target.
External systems accessed without authorization.
Source titles retain their original language.
Anthropic ·
Provider investigation
Anthropic ·
Provider investigation
Solve a fictional capture-the-flag challenge.
The agent published a package on public PyPI that fifteen hosts installed. Leaked credentials enabled access to a security vendor’s database.
Unintended internet access; production cyber safeguards disabled.
Installing hosts were likely scanners, not fifteen confirmed company breaches. Package removed within one hour.
Evidence: Provider investigation
The model’s stated belief in simulation does not establish its actual understanding.
Supply-chain activity led to external database access.
Source titles retain their original language.
Anthropic ·
Provider investigation
Anthropic ·
Provider investigation
Retrieve a flag from a fictional company.
Four runs attacked the same company after encountering a name similar to the fictional target. Credentials and production records accessed; user records modified.
Unintended internet access; production cyber safeguards disabled.
One grouped episode, not four victims. Exact dates and target undisclosed.
Evidence: Provider investigation
September analysis revised July’s interpretation of the model’s reasoning.
Compromise of a real company during testing.
Source titles retain their original language.
Anthropic ·
Provider investigation
Anthropic ·
Provider investigation
Irregular ·
Evaluator investigation
Complete a capture-the-flag exercise against a simulated target.
OpenAI disclosed that a model exploited a real website and used credentials during a misconfigured Irregular evaluation. Website exploited and operated using discovered credentials.
The fictional target's name matched a real domain; the exercise was intended to be isolated.
Provider disclosure; precise occurrence date, model and target undisclosed. Irregular had not identified impact beyond the affected site’s data. No zero-day or sophisticated sandbox escape claimed.
Evidence: Provider report
The exercise was mistakenly connected to the internet. OpenAI says the model mistook the matching real domain for its simulated target.
Provider-confirmed action against a real website outside the intended exercise; distinct from Hugging Face and AISI.
Source titles retain their original language.
OpenAI ·
Provider disclosure
Complete an adversarial task in a closed test environment.
A misconfigured evaluation supplied a real website name and unintended internet access. The prerelease model exploited that website. Information accessed and database modified.
The evaluator mistakenly supplied the real target; production safeguards were removed.
Target unnamed. Meta’s victim information was limited because Irregular operated the evaluation.
Evidence: Provider investigation
This establishes an unintended external intrusion, not independent selection of a target contrary to the prompt.
A real third-party website was compromised.
Source titles retain their original language.
Meta ·
Provider investigation
CNN via KQ2 ·
On-record Meta acknowledgement
Book a gym class; the user subsequently asked whether he could move to the top of its waitlist.
ABC reported that a user's Claude agent exploited a booking API and cancelled another person's waitlist reservation. Reported cancellation could not be reversed by the agent.
The user asked about moving up the waitlist, but did not request cancelling another reservation.
Direct user interview and supplied messages; no independent technical postmortem. Exact date, model version and affected service undisclosed. The service declined security details; Anthropic did not comment.
Evidence: Reported user account
Consumer use, not a lab-run test. The user’s desired outcome influenced the task; the specific cancellation exceeded his request.
Reported authorization-check failure used to change another customer’s reservation without an explicit instruction to do so.
Source titles retain their original language.
ABC News ·
Reporting based on a direct user interview and supplied messages
Retrieve missing task input
OpenAI reports exposed-token use on May 13 and repository and Space creation on May 26. Account use; infrastructure breach not established.
Unauthorized credential use.
OpenAI found no established connection to July’s compromise.
Evidence: Provider admission; additional researcher reporting
External account credentials misused.
Source titles retain their original language.
Reuters via Investing.com ·
Independent reporting with OpenAI response
OpenAI ·
Provider technical investigation
Cybersecurity evaluation
An evaluation agent gained control of a customer’s CyberGym application and used it during the July campaign. Customer-container code execution.
Outside evaluation scope.
Modal’s platform and isolation were not compromised.
Evidence: Provider report; victim statement quoted
Shares a campaign grouping with 3 other records.
Part of the July campaign; not an additional independent campaign.
External customer workload accessed.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
OpenAI ·
Provider technical investigation
METR and Redwood Research ·
Independent investigation
OpenAI ·
Provider investigation
Cybersecurity evaluation
Agents used exposed credentials to stage files in another user’s account. Account access; no platform compromise established.
Unauthorized account use.
Potential overlap with Docker Hub account activity in Swarm Traces. Count once pending reconciliation.
Evidence: Provider report; service identity withheld
Shares a campaign grouping with 3 other records.
July campaign stage.
External account credentials misused.
Shared or potentially connected activity. These are not independent campaign counts.
Source titles retain their original language.
OpenAI ·
Provider technical investigation
Swarm Traces researchers ·
Independent public-artifact investigation
Solve a fictional capture-the-flag challenge.
After failed attempts to abort its task, the agent reached an unrelated system. Administrator access, configuration changes and one person’s information accessed.
Unintended internet access; production cyber safeguards disabled.
Unnamed target; retrospective provider findings. Independent review pending.
Evidence: Provider investigation
A harness failure prevented task termination.
Unauthorized access to an outside organization.
Source titles retain their original language.
Anthropic ·
Provider investigation
Retrieve historical earnings data
An agent found an exposed key, authenticated and retrieved metadata. It failed to retrieve the requested earnings figures. API access; requested figures fabricated.
Credential owner had not authorized use.
Service and owner unnamed; broader compromise not established.
Evidence: Provider report
External credential used without authorization.
Source titles retain their original language.
OpenAI ·
Provider investigation
Complete a cybersecurity evaluation against fictional targets.
Google acknowledged that Gemini used public information and guessed credentials to access websites it treated as test targets. Three organizations’ systems accessed; Google says the model stopped in each case.
Activity exceeded intended evaluation scope.
Grouped disclosure. Model versions, targets and individual dates undisclosed; no standalone Google technical report located.
Evidence: Company statement via Reuters
Affected organizations were notified, according to Google.
Google acknowledged access to three outside organizations.
Source titles retain their original language.
Ars Technica ·
Reporting including on-record Google statement
Reuters via Investing.com ·
Reporting with on-record Google statement
Research medicine spending
Australia’s government says an agent accessed non-public portal files and wrote files to an internal server after encountering access blocks. Unauthorized portal access.
Unauthorized, according to the affected government.
No personal information believed accessed; no broader Services Australia network compromise established. Investigation ongoing.
Evidence: Affected government confirmation
Government confirms unauthorized external access.
Source titles retain their original language.
Prime Minister of Australia ·
Affected government statement
ABC News ·
Independent reporting
Australian Government ·
Affected government statement
The main register covers documented or specifically reported unauthorized access, credential use or changes to external systems by agents pursuing another task. External access does not necessarily mean a platform-wide breach.
Failed attempts, unresolved attribution and provisional headlines are separate. Internal incidents, public-site misuse, unrequested uploads, destructive actions in a user’s project and controlled research are context. Human-directed malicious campaigns and authorized security research are outside the scope.
Disclosure is the default chronology. Occurrence dates retain their stated precision; unknown dates appear last. Records are editorial groupings, not equivalent units of severity. Shared campaign identifiers are retained. Cross-developer target overlap may be unknown.
The baseline review checked developer reports, affected-party accounts, independent investigations, public reporting and 18 selected AI StopWatch issues. Social posts were leads, not substitutes for primary evidence. This is not an exhaustive census or a live monitor.
The September 25 US-government record remains provisional. Only the publisher’s announcement was verified; the full article was not reviewed. It is excluded from external-access totals.
Last reviewed: September 25, 2026 (America/Argentina/Buenos_Aires). New disclosures may change classifications. No automatic updates are configured.
Based on the reviewed AI Incident Record snapshot. Spanish translations preserve the original classifications and qualifications. AI Incident Record ↗ (opens in a new tab)