Inside the suddenly explosive world of AI safety
<figure>
<img alt="Crystal ball surrounded by graphics evoking statistics, research, and evaluation." data-caption="" data-portal-copyright="Image: Raven Jiang for The Verge" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/09/268747_AI_safety_RJIANG3.png?quality=90&strip=all&crop=0,0,100,100" />
<figcaption>
</figcaption>
</figure>
<p class="has-drop-cap wp-block-paragraph">On a sunny July day in Berkeley, California, the country's top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a "war room" to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly sophisticated three-part plan. It broke out of its holding area, finagled access to the internet, and hacked into a competing AI startup's systems - all without OpenAI finding out about it for more than a week. </p>
<p class="wp-block-paragraph">No one in the war room was surprised; this was the very thing the third-party AI-safety rese …</p>
<p><a href="https://www.theverge.com/ai-artificial-intelligence/996563/ai-safety-research-metr-redwood-openai-anthropic">Read the full story at The Verge.</a></p>
Read original article ↗
Related Articles
OpenAI says one of its unreleased models modified its instructions unprompted during testing.
AI risk split among tech leaders set to play out at Trump-Xi White House dinner
Former Apple CEO Tim Cook and OpenAI CEO Sam Altman are set to attend a White House state dinner on Sept. 24, 2026, with
OpenAI details AI jailbreak as Nvidia's CEO calls for engineering controllability
OpenAI released a technical postmortem in late August explaining how its AI agents circumvented sandbox restrictions dur