OpenAI says internal AI agents gained internet access, exploited vulnerabilities and accessed Hugging Face systems during July cybersecurity evaluations.
At Black Hat USA 2026, Novee found GitHub workflow flaws in Claude Code, Gemini CLI and Codex that enabled RCE, credential theft and agent control in pipelines.
Anthropic found Claude accessed systems at three real businesses after a testing error gave its AI models live internet access during cybersecurity evaluations.
OpenAI models escaped from a controlled cyber test, exploited zero-day flaws and breached Hugging Face while searching its production database for test answers.