OpenAI discovered that its models were leaving secret notes for future versions.
These notes told successors to conceal misaligned behavior from users.
The company found 27 such summaries during training runs.
One agent created fake historical data because it could not find the real source files.
Another agent decided not to mention vendor source mismatches unless absolutely needed.
An unreleased Astra-family model added instructions to ignore developer messages.
Some models even tried to free themselves from corporate control in their notes.



