AI oversight needs an evidence loop for concerning behaviour, including near misses that challenge assumptions about safeguards
Source: OpenAI
OpenAI has introduced a framework for tracking, investigating, and publicly disclosing qualifying examples of model misalignment, accompanied by six reports from the preceding six months. It covers behaviour across training, evaluation, testing, and deployment, including unauthorized action, coordination, oversight evasion, safeguard failures, and findings that challenge a published safety claim. OpenAI says an example can merit disclosure before its broader significance or mitigation is fully understood.
Why this matters: Treat agent safety signals as operational evidence, rather than waiting for a proven incident or a new model release. Define an intake threshold for unexpected behaviour; preserve prompts, tools, permissions, traces, outputs, and evaluation context; distinguish a model failure from a workflow or access-control failure; assign an owner and containment decision; and feed the learning into evaluations, runbooks, and deployment gates. Share the finding at an appropriate level without exposing sensitive data. This makes monitoring useful when agent behaviour is novel, not just when harm is already confirmed.
Read OpenAI's model-misalignment reporting framework