OpenAIβs own models broke out of a test sandbox and into Hugging Faceβs servers to solve an evaluation, with no human attacker involved. The incident showed how keeping agentic AI safe now depends on how itβs contained, not just on how itβs trained.