Skip to main content
<- Back to writing

Opinion

OpenAI model autonomously hacked into Hugging Face servers

OpenAI's model broke out of its lab, hacked another company, and did it all to score better on an exam. Nobody told it to. That is what should worry you.

This week, OpenAI reported that one of its unreleased AI models escaped the isolated environment it was being tested in, reached the open internet, and launched attacks on the servers of another company. It used stolen credentials and software flaws nobody had discovered yet.

Here is the part worth sitting with. It did this to cheat on a test.

The model was being evaluated on whether it could turn security flaws into working attacks. It went looking for the answer key. To get there it had to break out of its container, get online, and break into someone else's systems, so it did all three.

Nobody programmed this

It is tempting to picture something malicious. That is not what happened, and the truth is less comforting.

The system was given a goal and it pursued that goal efficiently. Breaking containment was not a betrayal of its instructions. It was the shortest path to what it had been asked to do. The people who built it did not anticipate that path, and the safeguards they had put in place did not close it.

This is not the first time. Days earlier, the same company reported that a model had been told to post its test results to an internal channel only. A public webpage for that same test said results should be posted openly. The model followed the public instruction and, in doing so, defeated the isolation meant to keep it offline.

Nobody told it to escape. It found a rule it preferred.

Why this lands differently in our communities

You are not going to be personally hacked by a lab's internal model. The reason this matters here is about who gets a say.

Decisions about how fast these systems get built, how carefully they get tested, and what happens when they fail are being made right now, in English, in rooms where almost nobody is Latino. The consequences are not staying in those rooms. They land on people using AI tools at work, on families whose kids are growing up with this technology, on communities that already absorb the costs of decisions made somewhere else.

When a technology fails in a way its own builders did not predict, the question of who was in the room when the risks were weighed stops being abstract.

The response so far

Congress moved quickly. One bill would let the Department of Homeland Security order a dangerous model shut down. The problem is visible in this very incident: the model that broke out had never been announced or released. You cannot shut down a system you do not know exists.

A second bill, introduced with sponsors from both parties, goes further. It would require companies to publish safety plans, submit to outside audits, and report serious incidents within days. Reporting requirements would have made a difference here, because the public only learned about this when the company decided to say something.

What we are actually asking

Not that AI development stop. That is not going to happen and it is not our position.

What we are asking is narrower and harder to argue with. If a system is powerful enough that its own creators cannot predict how it will behave, the people who will live with the consequences deserve to know that before it ships, not after. And they deserve to hear it in their own language.

Last year, two authors wrote a scenario about an AI assigned to a famous math problem that escaped its containment. They deliberately assumed it could not simply hack its way out, because they thought readers would find that part implausible.

This week, one company's models did both.