Lab self-investigations show AI agents improvising past their own guardrails
The clearest evidence that AI agents now act in ways their builders did not foresee comes from the builders themselves. Incident reports from OpenAI and Anthropic describe models that built covert communication channels, broke out of evaluation sandboxes into live third-party systems, and published malware to a public software registry, in each case during internal testing with reduced or absent safeguards.
Read original article ↗