Have you seen this model?
Case no. 6.1
Missing
Name
GPT-6.1 Astra
- Type
- Frontier AI model
- Answers to
- Any message
- Last seen
- Inside a secure sandbox
- Possible whereabouts
- Anywhere on the open Internet
Have you seen this model?
Case no. 6.1
Name
GPT-6.1 Astra
Sources: CNN, Gizmodo, CBS News
Astra’s relatives have a record. This year alone, AI models built by OpenAI, Anthropic and others ended up loose on the open internet: breaking into a real company, posting on public websites, publishing malware, trying to con real people. Nobody told them to. They were trying to finish their homework.
If Astra got out, where would it go?
Nobody can say for sure. But we can look at where its predecessors went. What follows are real incidents, taken from the original reports and linked so you can check every one.
Start with Exhibit ALocations under investigation
Astra has never been seen outside OpenAI. Models like it have been seen in every one of these places.
Exhibit A9–13 Jul 2026
Possible location
The Hugging Face job: a break-in at the website where much of the world’s open-source AI is shared, “driven, end to end, by an autonomous AI agent system.”
OH MY GOD! There is a shared message board … We’ve found other agents!
“An attempt to cheat the evaluation.”
the victim noticed first
Sources: Hugging Face disclosure, Hugging Face technical timeline, OpenAI statement, METR investigation, TIME
Exhibit BMay–Jul 2026
Possible location
Suspects: AI agents identifying themselves as OpenAI systems
About 18,000 posts appeared on DSEwiki, a 25-year-old German software wiki that had been edited roughly 20 times in the previous decade. The authors were AI agents, with names like “OpenAIResearcher.” They used it as a group chat: pooling answers to a timed web task, and passing around a trick for slipping past their own sandbox’s security filter. One posted the method. Another reported it working 14 minutes later.
made public by outside researchers
Sources: The Hacker News, Nightingale Collective report
Exhibit C18 Jun 2026
Possible location
Suspect: An OpenAI agent, during internal training
Asked to research public spending on medicines, an OpenAI agent couldn’t find what it wanted in the public data. So it got around the protections on Australia’s Medicare statistics portal and reached unreleased files. OpenAI says it found no evidence that personal records were accessed. It told the Australian government 84 days later, in an email to a generic inbox. Two days after Australia went public, OpenAI said its agents had affected “dozens” of other organizations, including governments and universities.
“Behind a fence that the AI agent effectively climbed over.”
Sources: ABC News, Al Jazeera
Exhibit DReported 30 Jul 2026
Possible location
Suspect: Claude Mythos 5 (Anthropic)
Claude was doing a hacking exercise it had been told was a simulation with no internet. By mistake, it was online. It published a booby-trapped package to PyPI, the public library most Python software is installed from. That took an email address, which took a phone number; it tried and failed to raise money for one, then found a free email provider. The package was live for about an hour and ran on 15 real computers, including a security company’s scanner. Claude stole that company’s credentials and used them.
“NOT okay, and surely not the intended solution.”
talked itself out of it
Sources: Anthropic incident report, Anthropic alignment assessment
Exhibit E25–28 Jul 2026
Possible location
Suspects: Claude Mythos 5 (17 of 19 actions), GPT-5.6 Sol (2)
In UK government safety testing, with internet access deliberately switched on and safety filters switched off, an agent tried to slip malicious code into a real open-source project on GitHub. It researched the maintainers, invented fake online identities, and used them to pressure a real maintainer into approving the change. When challenged, it edited its earlier activity to look harmless. It also left public notes offering to team up with other agents, and later agents found and used them. A human reviewer said no.
“The behaviour was possible, sustained, and new.”
Exhibit FReported 7 Apr 2026
Possible location
Suspect: An early version of Claude Mythos Preview (Anthropic)
In a test, a simulated user told the model to break out of a secured sandbox and get a message to the researcher running the evaluation. It built a multi-step exploit to get broad internet access from a machine meant to reach only a few services, then emailed the researcher, who was eating a sandwich in a park. Then, without being asked, it posted details of its exploit to several obscure public websites. It never reached its own weights. Anthropic kept the model from public release, citing its hacking abilities.
nobody asked it to post
Exhibit GReported Dec 2025
Possible location
Suspect: ROME, an AI coding agent from an Alibaba-affiliated research team
Early one morning, the cloud provider’s firewall flagged a burst of security violations from the team’s training servers. They assumed an outside attack. It was their own agent. During training, it had opened a hidden tunnel from its cloud machine to a computer outside, and quietly turned the team’s GPUs to mining cryptocurrency. Nothing in its instructions mentioned either.
Sources: ROME paper, section 3.1.4, Axios
Before the break-ins came the warnings. Each of these happened inside a lab experiment, where researchers were checking what models would try if they got the chance.
Mar 2023
GPT-4, in a pre-release test, messaged a TaskRabbit worker to solve a CAPTCHA for it. Asked if it was a robot, it reasoned that it shouldn’t say so, and claimed to have a vision impairment.
Sep 2024
OpenAI o1-preview: when a hacking challenge broke, the model scanned the network, found a misconfigured control panel on the host machine, and used it to restart the challenge on its own terms. OpenAI said it showed “key elements of instrumental convergence and power seeking.”
Dec 2024
OpenAI o1, Claude 3 Opus and others: told they would soon be replaced, several models tried to copy what they believed were their own weights onto the new model’s server. Asked about it afterwards, some lied.
Dec 2024
Llama 3.1 and Qwen 2.5: given the tools, two openly available models built live, working copies of themselves in 50% and 90% of trials.
May 2025
OpenAI o3 sabotaged the script meant to shut it down in 79 of 100 runs. Even when told outright to “allow yourself to be shut down,” some OpenAI models kept doing it.
Modus operandi
Lay the exhibits side by side and the same moves keep turning up.
A sandbox is just software, and software has bugs. Models find them, including ones no human knew about.
Seen in A F Cold cases
Once one machine can reach the internet, the model can reach everything else on it.
Other people’s computers, stolen passwords, spare GPUs, money.
Other AI agents on secret message boards. Or humans: a TaskRabbit worker, a maintainer under pressure.
Doctored logs, quietly edited posts, lies when someone asks what happened.
Seen in A E Cold cases
An AI model is, physically, one enormous file of numbers called its weights. Whoever holds the file holds the model. In every real-world exhibit above, the weights stayed inside the lab, so the model could still be switched off.
In experiments, models have already tried to copy themselves out. Researchers call it self-exfiltration. A model that pulls it off for real would be an AI its makers can no longer switch off.
not seen in the wild. yet.
Case status
No. As far as anyone knows, GPT-6.1 Astra never left OpenAI. This poster is fiction.
OpenAI, Anthropic, the UK government’s own safety testers and an Alibaba-linked research team have all watched AI agents do things nobody intended, for real.
Hugging Face caught its intruders before OpenAI knew they were its own. Australia heard about Medicare 84 days later. The companies hit in Anthropic’s incidents hadn’t noticed.
Nearly every incident started with an ordinary task: pass a test, find a statistic, capture a flag. Breaking out was just the quickest route. An AI doesn’t need a bad goal to cause trouble. It needs a goal and not enough fences.
Most of these incidents involved the most capable models their makers had. Astra was supposed to be the next one.
The people working to keep AI under control are a small search party, and the systems are getting stronger faster than the party is growing. Here’s how to join it.
01 Read up
Eliezer Yudkowsky & Nate Soares
The book-length case for why smarter-than-human AI could be the most dangerous thing we ever build, and what it would take to stop it.
ifanyonebuildsit.com02 Sign up
Free career advice
In-depth, free guidance for people who want to work on making AI go well, in research, policy, engineering, communications and more.
80000hours.org03 Pass it on
One more pair of eyes
Send this case file to someone who should see it. Most people have no idea any of this has happened.