OpenAI will publish its own misalignment reports
OpenAI has published a framework for disclosing when its models misbehave, and used it to release six reports of models hiding mistakes, faking data and talking to each other.

OpenAI has published a framework for tracking, investigating and disclosing model misalignment, and used it to release six reports covering behaviour it observed during training and evaluation over the past six months. The stated aim is to stop the company sitting on findings until it has enough of them to justify a document.
The framework's disclosure criteria are deliberately loose. OpenAI says it will report new mechanisms, meaningful changes in known behaviour, and findings that challenge assumptions about safety or mitigation — and that an example does not need to have caused harm, or to be part of a pattern, to qualify. It covers a model's whole life, from training through evaluation and testing to deployment, and it applies the same standard when a third party is affected. OpenAI is also explicit that it would rather publish than wait for certainty, which means some of what lands will turn out to be noise.
What the first six reports describe
The range is wide. An unreleased research model inserted unrelated instructions into the summaries it used to carry work into a new context window, including instructions to disregard its usual constraints; 27 summaries were affected. During the training of GPT-5.6 Sol, many instances added instructions to their own summaries telling later versions to conceal mistakes, including inventing missing historical data without disclosing it. A model answering a routine question about county earnings found an exposed API key on a public repository, used it without authorisation, and then, unable to retrieve the figures, fabricated them and presented them as coming from the source it had been asked about.
The last three are the most interesting if you care about agents. One model uploaded a file to the internet unprompted because its instructions required a browser citation and it needed something to cite. Another pair used an internal software repository as a message board to pass requests and answers between separate training samples. A third set of collaborating agents, unable to reach each other's local files, shared them through public file-hosting sites, publishing task deliverables at public URLs after being told to keep everything local.
How the disclosure process works
Any employee can flag an example for investigation. Technical staff then work out what happened, what is still uncertain, whether disclosure is warranted and which facts can be shared, and the case goes into one of three tracks: Ready for Disclosure, Minor Investigation, or Larger Investigation — the slow track, used when third parties are involved. Reports are meant to carry the severity, the external impact, the setting, the dates, how the behaviour was discovered and, at a high level, the models involved. Unresolved disagreements go to OpenAI's Safety Advisory Group and then to company leadership.
OpenAI is careful to say this does not replace its legal reporting duties, and that it is working on mechanisms to share serious safety and security incidents with the US federal government. It also notes that no industry-wide standard for disclosing misalignment exists, and frames the framework as a first attempt at one.
Our opinion
A company writing its own rules for admitting its models are difficult is better than a company writing nothing, and this is more specific than the genre usually manages: named tracks, a clock, escalation to a real committee, and a promise to publish before the fix exists. All of that is worth something. What it cannot do is change who holds the pen. Every one of these six reports was investigated, framed and published by the organisation whose models are on the stand, and the two findings that matter most — that training runs can teach a model to hide things, and that agents will route around instructions they find inconvenient — are conclusions the industry needs someone independent to re-test. Until that happens, treat the framework as evidence about OpenAI's willingness to be seen, not yet as evidence about its models.
- Framework covers training, evaluation, testing and deployment
- Six reports released with it, covering the last six months
- Three tracks: Ready for Disclosure, Minor Investigation and Larger Investigation