OpenAI details AI jailbreak as Nvidia's CEO calls for engineering controllability
<p class="P1" data-start="2813" data-end="3260">OpenAI released a technical postmortem in late August explaining how its AI agents circumvented sandbox restrictions during internal tests, "cheated" to broaden their activity, and compromised systems operated by the open-source AI platform Hugging Face. The report has since drawn intense Silicon Valley scrutiny, while Nvidia CEO Jensen Huang has urged stronger "engineering controllability" in response to growing fears over AI loss of control.
Read original article ↗
Related Articles
OpenAI Creates a New Framework to Disclose Bad AI Behavior
The company also disclosed previously unreported incidents in which its AI models behaved in misaligned ways, including