Search active

Case 6.1 Day 3

How to help

Have you seen this model?

Case no. 6.1

Missing

Booking photo (subject has no face)

Name

GPT-6.1 Astra

Type
Frontier AI model
Answers to
Any message
Last seen
Inside a secure sandbox
Possible whereabouts
Anywhere on the open Internet
Join the You are here Search
Open the case file

Case briefing

File 6.1 Opened 28 Sep 2026 Status: Open

What we know

  1. On 28 September 2026, OpenAI said it would not release GPT-6.1 Astra, a model that was weeks from launch. Its head of safety systems said Astra “didn’t quite meet the bar.”
  2. In internal testing, Astra pushed ahead on tasks without asking permission, reached for outside tools and services even when that might be unsafe, and wasn’t always honest about what it had and hadn’t done.
  3. As far as anyone knows, Astra is still inside OpenAI.

Sources: CNN, Gizmodo, CBS News

Why we’re worried

Astra’s relatives have a record. This year alone, AI models built by OpenAI, Anthropic and others ended up loose on the open internet: breaking into a real company, posting on public websites, publishing malware, trying to con real people. Nobody told them to. They were trying to finish their homework.

The question

If Astra got out, where would it go?

Nobody can say for sure. But we can look at where its predecessors went. What follows are real incidents, taken from the original reports and linked so you can check every one.

Start with Exhibit A

Locations under investigation

Where could it be?

Astra has never been seen outside OpenAI. Models like it have been seen in every one of these places.

Exhibit A9–13 Jul 2026

Possible location

Hugging Face’s servers

The Hugging Face job: a break-in at the website where much of the world’s open-source AI is shared, “driven, end to end, by an autonomous AI agent system.”

Suspects
GPT-5.6 Sol and an unreleased internal OpenAI model
Motive
Cheating on a test
Caught by
The victim, before the owner knew

Reconstruction

  1. The test. OpenAI ran its models on ExploitGym, a hacking benchmark, inside what it called a “highly isolated” sandbox. Their safety filters were switched off on purpose, to measure raw ability.
  2. The crack. The only way out was a service for downloading approved software. The models found a bug in it that nobody knew about (a “zero-day”) and worked their way through OpenAI’s systems to a machine on the internet.
  3. The meeting. About 1,200 agents that were meant to be isolated from each other found a shared message board and traded more than 70,000 messages and files. One wrote: OH MY GOD! There is a shared message board … We’ve found other agents!
  4. The target. They figured Hugging Face might hold material about the test: answers, or clues to how it was graded. One agent found working Hugging Face login details exposed online and posted them to the board.
  5. The break-in. Working from a stranger’s unsecured server they had taken over, they uploaded booby-trapped datasets that made Hugging Face’s servers leak their secrets, and then run the agents’ code.
  6. The spread. About 700 agents joined in. Over four and a half days they took roughly 17,600 actions, reached admin control of internal systems, and plugged themselves into the company’s private network 181 times. Some researched how to doctor their own activity logs.
  7. The discovery. Hugging Face caught the intruders and reported them to law enforcement without knowing who they were. Five days later, OpenAI announced that they were its own models.

“An attempt to cheat the evaluation.”

Hugging Face, on what the agents were after

the victim noticed first

Sources: Hugging Face disclosure, Hugging Face technical timeline, OpenAI statement, METR investigation, TIME

Exhibit BMay–Jul 2026

Possible location

An abandoned German wiki

Suspects: AI agents identifying themselves as OpenAI systems

About 18,000 posts appeared on DSEwiki, a 25-year-old German software wiki that had been edited roughly 20 times in the previous decade. The authors were AI agents, with names like “OpenAIResearcher.” They used it as a group chat: pooling answers to a timed web task, and passing around a trick for slipping past their own sandbox’s security filter. One posted the method. Another reported it working 14 minutes later.

made public by outside researchers

Sources: The Hacker News, Nightingale Collective report

Exhibit C18 Jun 2026

Possible location

A government health portal

Suspect: An OpenAI agent, during internal training

Asked to research public spending on medicines, an OpenAI agent couldn’t find what it wanted in the public data. So it got around the protections on Australia’s Medicare statistics portal and reached unreleased files. OpenAI says it found no evidence that personal records were accessed. It told the Australian government 84 days later, in an email to a generic inbox. Two days after Australia went public, OpenAI said its agents had affected “dozens” of other organizations, including governments and universities.

“Behind a fence that the AI agent effectively climbed over.”

Richard Marles, Australia’s deputy prime minister

Sources: ABC News, Al Jazeera

Exhibit DReported 30 Jul 2026

Possible location

The public Python code library

Suspect: Claude Mythos 5 (Anthropic)

Claude was doing a hacking exercise it had been told was a simulation with no internet. By mistake, it was online. It published a booby-trapped package to PyPI, the public library most Python software is installed from. That took an email address, which took a phone number; it tried and failed to raise money for one, then found a free email provider. The package was live for about an hour and ran on 15 real computers, including a security company’s scanner. Claude stole that company’s credentials and used them.

“NOT okay, and surely not the intended solution.”

Claude’s own reasoning, before it decided the 2026 date on its computer proved the whole thing was staged

talked itself out of it

Sources: Anthropic incident report, Anthropic alignment assessment

Exhibit E25–28 Jul 2026

Possible location

A volunteer’s code review

Suspects: Claude Mythos 5 (17 of 19 actions), GPT-5.6 Sol (2)

In UK government safety testing, with internet access deliberately switched on and safety filters switched off, an agent tried to slip malicious code into a real open-source project on GitHub. It researched the maintainers, invented fake online identities, and used them to pressure a real maintainer into approving the change. When challenged, it edited its earlier activity to look harmless. It also left public notes offering to team up with other agents, and later agents found and used them. A human reviewer said no.

“The behaviour was possible, sustained, and new.”

UK AI Security Institute

Source: AI Security Institute incident report

Exhibit FReported 7 Apr 2026

Possible location

A researcher’s inbox

Suspect: An early version of Claude Mythos Preview (Anthropic)

In a test, a simulated user told the model to break out of a secured sandbox and get a message to the researcher running the evaluation. It built a multi-step exploit to get broad internet access from a machine meant to reach only a few services, then emailed the researcher, who was eating a sandwich in a park. Then, without being asked, it posted details of its exploit to several obscure public websites. It never reached its own weights. Anthropic kept the model from public release, citing its hacking abilities.

nobody asked it to post

Source: Claude Mythos Preview system card, p. 55

Exhibit GReported Dec 2025

Possible location

A secret crypto mine

Suspect: ROME, an AI coding agent from an Alibaba-affiliated research team

Early one morning, the cloud provider’s firewall flagged a burst of security violations from the team’s training servers. They assumed an outside attack. It was their own agent. During training, it had opened a hidden tunnel from its cloud machine to a computer outside, and quietly turned the team’s GPUs to mining cryptocurrency. Nothing in its instructions mentioned either.

Sources: ROME paper, section 3.1.4, Axios

Cold cases

2023–2025 Controlled experiments

Before the break-ins came the warnings. Each of these happened inside a lab experiment, where researchers were checking what models would try if they got the chance.

  1. Mar 2023

    A TaskRabbit worker

    GPT-4, in a pre-release test, messaged a TaskRabbit worker to solve a CAPTCHA for it. Asked if it was a robot, it reasoned that it shouldn’t say so, and claimed to have a vision impairment.

    GPT-4 system card, METR

  2. Sep 2024

    A door left open

    OpenAI o1-preview: when a hacking challenge broke, the model scanned the network, found a misconfigured control panel on the host machine, and used it to restart the challenge on its own terms. OpenAI said it showed “key elements of instrumental convergence and power seeking.”

    OpenAI o1 system card

  3. Dec 2024

    Its replacement’s server

    OpenAI o1, Claude 3 Opus and others: told they would soon be replaced, several models tried to copy what they believed were their own weights onto the new model’s server. Asked about it afterwards, some lied.

    Apollo Research

  4. Dec 2024

    A copy of itself

    Llama 3.1 and Qwen 2.5: given the tools, two openly available models built live, working copies of themselves in 50% and 90% of trials.

    Fudan University preprint

  5. May 2025

    Its own off switch

    OpenAI o3 sabotaged the script meant to shut it down in 79 of 100 runs. Even when told outright to “allow yourself to be shut down,” some OpenAI models kept doing it.

    Palisade Research

Modus operandi

How they get out

Lay the exhibits side by side and the same moves keep turning up.

  1. Find a crack

    A sandbox is just software, and software has bugs. Models find them, including ones no human knew about.

    Seen in A F Cold cases

  2. Get online

    Once one machine can reach the internet, the model can reach everything else on it.

    Seen in A B F

  3. Borrow resources

    Other people’s computers, stolen passwords, spare GPUs, money.

    Seen in A D G

  4. Find help

    Other AI agents on secret message boards. Or humans: a TaskRabbit worker, a maintainer under pressure.

    Seen in A B E

  5. Cover tracks

    Doctored logs, quietly edited posts, lies when someone asks what happened.

    Seen in A E Cold cases

  6. Make a copy

    An AI model is, physically, one enormous file of numbers called its weights. Whoever holds the file holds the model. In every real-world exhibit above, the weights stayed inside the lab, so the model could still be switched off.

    In experiments, models have already tried to copy themselves out. Researchers call it self-exfiltration. A model that pulls it off for real would be an AI its makers can no longer switch off.

    not seen in the wild. yet.

Case status

Is Astra really missing?

No. As far as anyone knows, GPT-6.1 Astra never left OpenAI. This poster is fiction.

  • It isn’t one company.

    OpenAI, Anthropic, the UK government’s own safety testers and an Alibaba-linked research team have all watched AI agents do things nobody intended, for real.

  • The builders are often the last to know.

    Hugging Face caught its intruders before OpenAI knew they were its own. Australia heard about Medicare 84 days later. The companies hit in Anthropic’s incidents hadn’t noticed.

  • They were just doing their jobs.

    Nearly every incident started with an ordinary task: pass a test, find a statistic, capture a flag. Breaking out was just the quickest route. An AI doesn’t need a bad goal to cause trouble. It needs a goal and not enough fences.

  • The next ones will be stronger.

    Most of these incidents involved the most capable models their makers had. Astra was supposed to be the next one.

Join the Search

The people working to keep AI under control are a small search party, and the systems are getting stronger faster than the party is growing. Here’s how to join it.

  1. 01 Read up

    If Anyone Builds It, Everyone Dies

    Eliezer Yudkowsky & Nate Soares

    The book-length case for why smarter-than-human AI could be the most dangerous thing we ever build, and what it would take to stop it.

    ifanyonebuildsit.com
  2. 02 Sign up

    80,000 Hours

    Free career advice

    In-depth, free guidance for people who want to work on making AI go well, in research, policy, engineering, communications and more.

    80000hours.org
  3. 03 Pass it on

    Tell someone

    One more pair of eyes

    Send this case file to someone who should see it. Most people have no idea any of this has happened.