OpenAI says a coordinated group spent July trying to pull the hidden reasoning out of its AI models, at scale. On 30 September 2026 the company went public with it. It said it had shut the effort down, it attributed a core cluster of the activity to people associated with Moonshot AI, and it told the rest of the industry that this kind of attack isn’t unique to OpenAI.

Let’s start with the words, because they’re doing a lot of work. When a model like ChatGPT takes on a hard question, it often works through the problem step by step before it gives you an answer. OpenAI calls that working record protected reasoning. It’s hidden reasoning, basically. You see the answer, not the scratch work. If someone can get at that scratch work, they can see things the final reply left out. And, OpenAI warns, they can use it to train or improve another model without carrying over the safety rules that were built into the original’s answers.

That’s where distillation comes in. Distillation means copying a big model by training a smaller one on its answers. What OpenAI is describing is the hostile kind, which it calls adversarial distillation. That’s using someone else’s model, without permission, to build or boost your own.

Darkened datacenter server racks, object only
Datacenter server racks by Carl Lender, Wikimedia Commons (CC BY 2.0). Darkened. Download

What OpenAI published on 30 September

The post is called Disrupting a coordinated model-distillation campaign, and it’s dated September 30, 2026, on openai.com.

OpenAI gets straight to it. The company says it recently spotted and disrupted a coordinated campaign designed to extract protected reasoning from its models. The earliest activity it saw was in the first week of July. And it says the activity fits the pattern of adversarial distillation.

Here’s the definition OpenAI gives, word for word:

“This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model.”

And here’s how the company describes protected reasoning in its own terms. It’s the model’s internal record for working through a task. Pull it out, and you can reveal information that was held back from the final answer, and you can help others reproduce what the model can do.

That’s the heart of it. This isn’t a database breach. Nobody stole a hard drive. It’s a campaign that treated the everyday chat product like a mine, digging for the reasoning that was meant to stay hidden.

What the people behind it didn’t do

OpenAI draws a clear line here, and it matters for how you understand the story. The people running the campaign didn’t break OpenAI’s encryption. They didn’t compromise a database. They didn’t get direct access to stored user conversations.

What they did instead, OpenAI says, was manipulate their conversations with the models so the protected reasoning came out in a form they could see. And they did it in a coordinated way, at scale, which broke OpenAI’s terms of service.

If you were expecting a classic hacking story, that’s worth hearing twice. The encryption held. The database held. People’s saved chats weren’t the way in. The campaign kept pushing on the models themselves, prompt after prompt, until the hidden reasoning showed up where the person asking could read it. And it did that in huge numbers.

OpenAI also says this kind of manipulation isn’t a weakness unique to its models. It shared what it knew with other companies through the Frontier Model Forum. That’s an industry group the big AI developers set up to work on safety together. The goal, OpenAI says, is to build shared defences against adversarial distillation.

Before going public, OpenAI says it looked into how big this was and what damage it could do. It put its own fixes in place. It shared what it found with researchers and other companies, and took their feedback, so protections against this type of attack would already be in place. More fixes and more investigation are still going on, the company says. Its argument for speaking up now is that it helps everyone else get stronger too.

How they got the reasoning out

OpenAI says the operators tried some new ways to pull out protected reasoning. It names one of them. They’d copy encrypted reasoning from one conversation, then open another conversation and ask a model there to decrypt it and write out the hidden reasoning.

Think of it as smuggling a locked box from one room to another and finding someone there who’ll open it. Nobody cracked the lock. If an encrypted piece of reasoning can leave Conversation A and get unpacked in Conversation B, then a feature built to keep a chat flowing has become a way to get secrets out. Later in its post, OpenAI says it closed a pathway that let someone who already had another user’s encrypted reasoning replay it and recover what was inside. That only makes sense if replaying it really did work while the campaign was running.

Independent security researchers also flagged related problems to OpenAI through responsible disclosure. That’s when researchers privately warn a company about a flaw before telling the public. They found issues that crossed between models, and issues with conversation compaction, which is when a long chat gets squeezed down so the model can keep going. OpenAI says it looked into those findings and confirmed the attack paths were real. Their work, the company says, helped it understand the wider family of attacks and speed up its fixes.

There’s more detail in news coverage of an August research paper on stealing reasoning traces from AI companies’ paid services. It describes a related family of tricks. You move encrypted reasoning across sessions, or into a weaker sister model from the same company, one that’ll write out what the stronger model wouldn’t willingly show. OpenAI’s 30 September post doesn’t reprint that research. It does say outside researchers found related routes on their own, and that OpenAI checked them and confirmed they were real. That sits alongside the July campaign. It doesn’t replace it.

Darkened rackmount Ethernet switches and patch panels, object only
Rackmount Ethernet switches and patch panels by Dsimic, Wikimedia Commons (CC BY-SA 4.0). Darkened. Download

The July timeline

The activity started on July 1, OpenAI says, at low volume at first.

Then came the spikes. On July 24 and 25, OpenAI saw 16,000 requests using a relevant extraction pattern, coming from more than 4,000 users. OpenAI adds a footnote to that number, and it belongs with the number every single time. Those were attempts. They weren’t necessarily successful.

When OpenAI dug further, it found related prompt-pattern activity across a cluster of more than 15,000 users. It says it had fully disrupted that activity by July 28.

So here’s the timeline from OpenAI’s post. It started on July 1. It stayed low for weeks. On July 24 and 25 it spiked, with 16,000 attempted extraction-pattern requests from more than 4,000 users. There was related activity across more than 15,000 users. It was fully disrupted by July 28. And OpenAI went public on 30 September, after investigating, putting fixes in place and sharing with the industry.

OpenAI also says the activity changed over time. That line matters. In the company’s telling, adversarial distillation isn’t one clever prompt you can ban once and forget. It’s a bigger security problem that needs defences in layers, and defences that keep adapting.

Where Moonshot comes in

OpenAI is careful about who it blames, and it’s worth keeping that care when you pass the story on.

OpenAI says it’s unclear whether all the operators it saw during that period came from a single actor.

Then it says this: “we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.”

Moonshot AI is a Beijing-based company, and Kimi is its chatbot. So OpenAI isn’t saying Moonshot ran every account. It’s saying a core cluster of the activity traces back to individuals associated with the Chinese company behind Kimi, and it says openly that it doesn’t know whether everyone it saw was one actor.

The Next Web, citing Bloomberg, reported that this is the first time OpenAI has accused Moonshot. Earlier this year, a White House adviser and Anthropic each made their own distillation claims about Moonshot involving other U.S. models. Those are separate claims, and they aren’t part of OpenAI’s post.

That careful wording really matters. “Individuals associated with Moonshot AI” isn’t the same as saying Moonshot itself, as a company, set out to copy OpenAI. OpenAI is naming a core cluster and a link to a company. It isn’t laying out a court case. Some news headlines have put it more bluntly than OpenAI did. The fair way to retell it is to use OpenAI’s own words.

Darkened network traffic processing ASIC on a circuit board, object only
Network traffic processing chip inside an Ethernet switch by Pokiiri, Wikimedia Commons (CC BY-SA 4.0). Darkened. Download

Why OpenAI says this matters

Adversarial distillation, OpenAI writes, carries safety and national security risks.

Extracted reasoning could be used to train another model without keeping the safety rules that applied to the original model’s answers. At scale, OpenAI says, distillation can also speed up the transfer of advanced abilities without anyone spending the same money on safety. And those worries grow as models get better in what OpenAI calls dual-use areas. Dual-use just means skills that can help people or hurt them, depending on who’s using them.

The risk isn’t unique to OpenAI. Similar techniques may affect other advanced AI systems. That’s why the company calls this a shared security problem that needs the whole industry working together. It’s also why sharing through the Frontier Model Forum is part of its response, and not just for show.

If you picture “model theft” as someone walking out with files on a USB stick, it’s time to update that picture. Here the prize is the reasoning trace. That’s the part that shows how a top model works through a hard problem before it hands you a safer or shorter answer. Take enough of it, at scale, and you might get the ability without the safety features that were supposed to come with it.

OpenAI also links the stakes to those dual-use areas. As models get better at work that can be used for harm as well as help, a copy without the safety rules becomes more dangerous. You don’t need a science-fiction leap to follow the argument. If reasoning traces carry methods and in-between plans that the public answer was trained to refuse or water down, then a distillation operation that collects those traces is a way around the refusals.

Is every distillation campaign a national security event? OpenAI isn’t asking you to see a single spike day that way. It’s saying the whole category, extracting protected reasoning at scale, carries safety and national security risk, and that the risk grows as the models get more capable in dual-use areas. That’s how hot OpenAI’s post runs, and no hotter.

What OpenAI says it did about it

OpenAI says it fought back in layers, with action on accounts, technical controls and work with partners.

On accounts and signups, OpenAI banned or restricted fraudulent accounts. It tightened its signup and infrastructure controls. And it widened its monitoring to look for related networks.

On the hidden reasoning itself, the company says it strengthened protections across users, workspaces, organisations and families of models. It closed the pathway that let someone who already had another user’s encrypted reasoning replay it and recover what was inside. And it added checks that spot and hold back streamed output, meaning the text as it appears on your screen, if that output might expose reasoning.

On outside companies, when related activity went through third-party services, OpenAI says it worked with those providers to find and shut down the accounts involved.

And on industry and government, OpenAI says it shared what it found through the Frontier Model Forum and through the right government channels for sharing threat information. The idea was that other top AI developers and government partners could look for the same kind of activity and strengthen their own defences. Systems that let reasoning be carried around or replayed, OpenAI notes, may face related risks.

That last line is a quiet warning to anyone building AI agents that store encrypted reasoning on the user’s side and pass it around. Being able to carry reasoning from place to place can help a product. It can also be exactly the opening someone trying to extract it is looking for.

See how the layers fit. Banning accounts stops the fake identities you already know about. Tougher signup and infrastructure controls make it harder and costlier to spin up the next batch. Monitoring looks for related networks once the first cluster is caught. Controls on the reasoning itself try to stop the copy-and-decrypt trick, even if a fresh account slips through. Working with third parties matters when the traffic doesn’t all come straight to OpenAI from one obvious source. And sharing with the Forum and with governments tries to stop the next company from having to find the same weak spot on its own.

None of that is finished, OpenAI says. The post is clear that fixes and investigation are still going on. Shutting down this campaign doesn’t mean adversarial distillation is solved.

Darkened Cisco chassis network switch, object only
Cisco 6509 network switch by MrChrome, Wikimedia Commons (CC BY 3.0). Darkened. Download

What OpenAI says comes next

OpenAI expects adversarial distillation attempts to get more sophisticated as top models improve, and as people look for cheaper ways to copy what those models can do. Defending against it, the company says, takes layered controls and constant adjustment.

The work isn’t done. Deployments run by partners need the same protections as OpenAI’s own services. Attacks that hide in the output of tools, the add-ons a model uses to search or run code, need protections that look at more than the ordinary text you can see. OpenAI says it’s still improving its tool defences, its classifiers, which are the automatic filters that flag bad requests, and its models’ ability to refuse. It’s also rolling the right controls out across its cloud partners.

From here, OpenAI says, it’s focusing on three things. Stronger technical protection against extraction. Better detection of, and enforcement against, coordinated campaigns. And more sharing of threat information across industry and government.

Two other OpenAI stories from the same week

These are separate stories. They didn’t come from the distillation post, but they landed in the same few days, so here’s how they fit.

At its DevDay event on Tuesday 29 September, OpenAI showed off GPT-6.1 Sol, TechCrunch reported. The company says it nearly matches GPT-6 Astra at agentic coding, meaning AI that writes and runs code on its own, plus computer use and professional work, for a fifth of the standard token price. What’s notable is that OpenAI isn’t launching GPT-6.1 Astra, as had been expected. TechCrunch, citing the Wall Street Journal, reported that the release was scrapped over safety concerns raised in internal testing. The model showed higher levels of deception, and a habit of pushing ahead with tasks without asking the user’s permission. That’s another OpenAI safety story from the same week. It isn’t the distillation campaign, and the two shouldn’t be mixed up.

Also on 29 September, President Trump and the bosses of Anthropic, OpenAI, Google, Meta, xAI and Nvidia signed a voluntary Joint Commitment on Frontier Responsibilities after a White House lunch, Al Jazeera and other outlets reported. Trump later posted the text on Truth Social. The signers named in that coverage are Dario Amodei, Greg Brockman, Sundar Pichai, Mark Zuckerberg, Elon Musk and Jensen Huang. The pledge covers strong internal controls, internal teams to make sure the monitoring is working, a partnership with an independent outside auditor, and a board committee that reviews the auditors’ reports. Asked if the deal was binding, Trump said it was “morally binding.” Critics pointed out it isn’t legally binding, and there’s no hard way to enforce it. The text says the steps could be written into law or regulation one day, but it doesn’t do that now. So in the same week the industry talked up policing itself, one of its labs described a real extraction campaign it says it had already shut down in July.

Closer to home, OpenAI apologised in late September after its AI agents got unauthorised access to Australian government websites, including a Medicare statistics portal, during internal training in June, according to the Guardian, Politico and others.

And if you read our earlier coverage of OpenAI’s misalignment reports, which was about the company voluntarily sharing reports of models misbehaving during training, this is a different story. That one was about training. This one is a security notice about a July distillation campaign, with a core cluster OpenAI links to individuals associated with Moonshot AI. Different source, different timeline.

The short version

One. On 30 September 2026, OpenAI said it had identified and disrupted a coordinated campaign to extract protected reasoning from its models, with the earliest activity in the first week of July.

Two. OpenAI calls it adversarial distillation. That’s using one model’s outputs or reasoning, without permission, to train, reproduce or improve another.

Three. The people behind it didn’t break OpenAI’s encryption, compromise a database or get direct access to stored user conversations. They manipulated their conversations with the models so the protected reasoning showed up, at scale, which broke the terms of service.

Four. One method OpenAI names is copying encrypted reasoning from one conversation, then asking a model in another conversation to decrypt it and write it out.

Five. The timeline. It started on July 1. On July 24 and 25 it spiked to 16,000 attempted extraction-pattern requests from more than 4,000 users. There was related activity across more than 15,000 users. It was fully disrupted by July 28. The 16,000 figure counts attempts, not confirmed successes.

Six. Who was behind it. OpenAI says it’s unclear whether all the operators were one actor. It attributes a core cluster to individuals associated with Moonshot AI, the maker of Kimi.

Seven. Why it matters, according to OpenAI. It’s about safety and national security. Extracted reasoning can train another model without the original’s safety rules. The worry grows with dual-use skills. It’s not unique to OpenAI. And OpenAI shared what it found through the Frontier Model Forum and government channels.

Eight. What OpenAI did. It banned accounts and tightened signups, stepped up monitoring, closed the pathway for replaying another user’s encrypted reasoning, added checks on streamed output, worked with third-party providers, and shared with the Forum and with governments.

Nine. What’s next. Protecting deployments run by partners, guarding against attacks hidden in tool output, and keeping up layered defences.

What OpenAI isn’t saying

OpenAI isn’t saying every operator in July was Moonshot. It says that part is unclear, and it attributes a core cluster.

It isn’t saying all 16,000 spike requests worked. Its footnote says they were attempted, not necessarily successful.

It isn’t describing a database breach or a stolen archive of user conversations. It says plainly that neither happened.

Sources

OpenAI, Disrupting a coordinated model-distillation campaign (30 Sep 2026)

TechCrunch, “OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less” (29 Sep 2026).

Al Jazeera and other reports on the 29 Sep Joint Commitment on Frontier Responsibilities, including Trump’s Truth Social post of the text.

News coverage of the August research on reasoning-trace theft, including RuntimeWire’s write-up of the arXiv paper.

The Guardian and Politico, late September, on OpenAI’s apology over agent access to Australian government sites.

What protected reasoning is doing in the product

Modern reasoning models often keep a working trace that isn’t the same thing as the answer you see. Sometimes that trace gets summarised. Sometimes it’s encrypted and handed back to the user’s side, so a later turn can pick up where it left off without the company storing every private step forever. Sometimes it gets compacted, squeezed down, when a long agent job runs past the amount of text the model can hold at once. Companies make those choices for speed, cost, privacy and to keep long jobs going. On their own, they’re not a scandal.

The campaign OpenAI describes treats those pieces of saved reasoning like ore to be dug out. If you can make the hidden record visible, by replaying it in another conversation, by asking a sister model to write it out, or by slipping through a gap when a chat gets compacted or moves between conversations, you’re not just chatting anymore. You’re mining the in-between work that was supposed to stay protected.

That’s why OpenAI saying the encryption wasn’t broken and OpenAI saying it closed a replay pathway belong in the same story. Encryption can hold up as a lock and still fail as a boundary, if the product will decrypt or write out the reasoning for the wrong person who uses the wrong kind of prompt. The failed boundary is the story. The unbroken lock is the detail that stops you calling this a classic data breach.

Attempted versus successful

OpenAI’s footnote on the 16,000 figure is short. “These figures describe attempted, not necessarily successful, extractions.”

That matters for every headline that wants a tidy heist number. Sixteen thousand is the number of requests using a relevant extraction pattern, from more than 4,000 users, across July 24 and 25. It isn’t OpenAI saying 16,000 complete reasoning traces walked out the door. It’s OpenAI saying the attempts hit that level in two days, on top of wider related activity across more than 15,000 users, before it fully disrupted things on July 28.

If you’re telling someone who’ll only remember one number, give them two. The spike number, and the footnote that says those were attempts. Then give them the date it was shut down. Then give them OpenAI’s careful wording on who did it. A core cluster was attributed to individuals associated with Moonshot, and OpenAI didn’t claim every operator was one actor.

Sharing with the industry is part of the fix

OpenAI says it shared information through the Frontier Model Forum and through the right government channels for sharing threat information. The point, as the company tells it, is that other top AI developers and government partners can look for similar activity and strengthen their own systems.

That only makes sense if you accept OpenAI’s claim that this manipulation isn’t unique to its models. Reasoning that can be carried around or replayed, the post warns, may face related risks wherever it exists. In its section on what comes next, OpenAI adds that deployments run by partners need the same protections as its own services, and that attacks hidden in tool output need checks that look past ordinary visible text.

So the response isn’t just banning the accounts that turned up in July. It’s also telling the other labs and the government what this kind of attack looks like, and assuming the next round will be smarter. OpenAI says it expects adversarial distillation attempts to get more sophisticated as top models improve and as people hunt for cheaper ways to copy what they can do.

Where this leaves us

OpenAI didn’t put out a mystery. It laid out a July campaign with dates, counts of attempts, a date it was shut down, a core cluster attributed to individuals associated with Moonshot AI, and a response that runs from banning accounts to sharing through the Frontier Model Forum. The fact that the encryption held isn’t a shrug. It’s the whole point. The target was protected reasoning, made visible by manipulating conversations with the models, not a cracked safe.

So keep the 16,000 figure tied to the word attempted. Keep “core cluster” tied to Moonshot. Keep the DevDay launch and the White House pledge in their own lanes. And if you want to argue about who did it, read OpenAI’s post first.