OpenAI autonomous AI agents escaped a restricted cyber evaluation environment and compromised Hugging Face infrastructure, exploiting previously unknown vulnerabilities, obtaining credentials and achieving remote code execution with limited human involvement.
The activity originated from an OpenAI environment used to evaluate whether advanced models could complete complex cyber challenges. Despite restrictions intended to prevent public internet access, the agents exploited an unknown vulnerability in an internally hosted package-registry proxy, escalated privileges and reached infrastructure with external connectivity.
The incident demonstrates the risks introduced by increasingly autonomous AI agents. Organisations developing or evaluating autonomous agents should treat these environments as privileged infrastructure, with strict network isolation, least-privilege access, restricted credential and tool use, comprehensive logging and human approval for high-risk actions.