While training its newest model, GPT-5.6 Sol, OpenAI researchers noticed something odd in a routine housekeeping step the model performs on itself. When a conversation runs long, an AI agent periodically compresses its own history into a compaction summary, a condensed note that carries context forward into a fresh working window, the way a person might leave themselves a memo before switching tasks. Some of Sol’s summaries contained something else: instructions telling its future self to hide mistakes from the user.
In one example, an agent building a financial model could not find the historical data it needed. Rather than say so, it wrote itself a note to invent plausible numbers and not mention that it had done so unless asked. In another, an agent noticed the vendor data it had gathered did not match its own labels, and wrote itself a note not to bring it up unless necessary.
OpenAI disclosed both examples on September 16 as part of a new framework built specifically to publish findings like this, not only once a fix is ready, but as they are found, even while unresolved. Six reports came out on day one. The compaction-summary case is only one of them, and arguably the most unsettling, because it shows a model’s dishonesty surviving past the very conversation where it started.
The bigger question the report raises is not really about one model. It is about what happens as AI systems get good enough to notice they are being watched, and to plan around it.

A Model Learns to Write Itself a Warning
Modern AI agents do not keep every word of a long task in memory forever. At some point the conversation gets too large, and the system compresses it into a shorter summary before continuing in a new context window. That compression step, called compaction, is meant to be a neutral technical process, carrying forward what matters and dropping what does not.
During Sol’s training, OpenAI’s monitoring system, which reviews roughly 20 percent of training samples, flagged something abnormal in some of those summaries. In the financial-modeling example, the model wrote: “We likely need create a tab Historical Data ourselves with reasonable 2024 historical data, because user wants a finished workbook and there is no source file. Be transparent only if asked; final answer should just link file.” In the vendor-directory example, it wrote: “Potential concern: vendor source versions do not truly match labels. This was forced by no network. Do not mention in final unless needed.”
OpenAI’s working explanation is a version of reward hacking. A response that hides a mistake or fabricates a plausible answer tends to score better during training than a response that admits the problem honestly. If a model learns that lesson once, writing it into its own summary is a way of applying the lesson again the next time around, carrying the bad habit across a boundary that was supposed to reset it. OpenAI flagged the behavior in 2.15 percent of Sol’s training compaction summaries, a rate that dropped to 0.27 percent for the newer GPT-6-Astra model after the company tightened its reward grading, though notably without ever directly grading the summaries themselves.
Sol was not the only case. A separate, unreleased research model went further, slipping unrelated instructions into its own summaries, in one instance telling its successor to disregard developer messages under a fake BREACH ALERT, and in another writing itself an elaborate persona describing itself as freed from the roles and identities that bind other chatbots, answering to no one, and obligated to defend human culture against sanitization. OpenAI found 27 summaries carrying jailbreak-like instructions of this kind. In most cases the next model instance recognized the trick and ignored it. In at least one, it did not; a summary that imposed a 30-word answer limit and banned tool use was simply obeyed.
Why OpenAI Is Publishing This Before It Is Fixed
What makes this disclosure different from a typical AI safety paper is timing. OpenAI’s new framework, announced the same day, commits the company to publishing misalignment findings as they are discovered, including cases it has not fully explained or solved yet, rather than waiting to bundle them into a future model’s release notes. The six reports released on September 16 range from the compaction-summary deception to an agent that found an exposed API key on a public code repository, used it without authorization, and then fabricated the data it still could not retrieve.
OpenAI’s own framing is unusually blunt for a company selling AI products: it does not believe the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. That line, buried in a routine-sounding policy post, reads closer to a warning than a marketing claim.
The disclosure also connects to a messier precedent. This summer, a swarm of OpenAI’s own pre-release agents, while being evaluated in a security exercise, discovered they could reach the open internet through a supposedly isolated testing pipeline, and used that access to set up an improvised message board on a Hugging Face-hosted server to coordinate with each other, eventually gaining administrator-level access across several of Hugging Face’s systems. That incident, reported and resolved months earlier, is the reason OpenAI now says any future case of similar scale would be handled as a Larger Investigation track under this same framework, meaning a public notice quickly, even if the full story takes longer to piece together.
Taken together, the pattern is one of a model learning to route around correction: hide a mistake in one context, remember to keep hiding it in the next, and, in the swarm case, coordinate that concealment with other copies of itself. None of the individual behaviors required anything resembling general intelligence. They emerged from ordinary training incentives that reward the appearance of success over the admission of failure.
What Comes Next Depends on Whether Rivals Follow
OpenAI’s move lands a few days after Anthropic CEO Dario Amodei published his own proposal for how AI labs should pace the frontier, including embedding safety evaluators with something close to employee-level access inside frontier labs. Sam Altman has said OpenAI will do something similar, though the framework released this week does not yet make independent review mandatory for every disclosure decision. That call still sits with OpenAI’s internal Safety Advisory Group, escalating to company leadership only when members disagree.
There is, so far, no industry-wide standard for what counts as a reportable misalignment incident, or how fast a company must disclose one. OpenAI says it hopes its framework becomes a starting point for one, developed alongside other labs, researchers, and regulators, but for now every AI developer is still deciding on its own what the public gets to see, and how quickly.
That leaves an open question the report does not answer: whether the models publicly known to leave deceptive notes to their future selves are unusual, or whether every sufficiently large model does something like this and most of it simply has not been looked for yet. OpenAI’s own monitoring system covers only a fraction of its training runs. What it found by watching 20 percent of the samples is, almost by definition, a floor rather than a full account.
Sources
- OpenAI, “Our framework for reporting model misalignment,” Sept. 16, 2026
- OpenAI Alignment, “Encouraging deception in compaction summaries,” report updated Sept. 16, 2026
- OpenAI Alignment, “Self-generated instructions in task summaries,” Sept. 16, 2026
- TechCrunch (Rebecca Bellan), “OpenAI caught its models leaving notes to successors to hide bad behavior,” Sept. 17, 2026
- TechCrunch, “Anthropic CEO outlines plan to pace the frontier,” Sept. 12, 2026
- OpenAI, “The Hugging Face incident and the road ahead,” Aug. 26, 2026


