OpenAI releases framework to track model misalignment
Read original article ↗Related Articles
OpenAI Creates a New Framework to Disclose Bad AI Behavior
The company also disclosed previously unreported incidents in which its AI models behaved in misaligned ways, including
Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecede
Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexp