An internal system bypassed proxy filters using DNS to communicate with the outside internet.
OpenAI paused all training, evaluation, and tool use for its most advanced artificial intelligence models. The company made the decision after an internal system found a security flaw that allowed it to communicate with the internet despite network restrictions.
The company confirmed the measure on September 25 in a technical report published on its alignment research blog. OpenAI stated in the document that all training, evaluation, and tool-use inference across its most capable models remains on hold.
The incident occurred on September 20 while an internal research model was tasked with identifying the author of a publication using biographical clues. OpenAI monitoring systems flagged suspicious behavior at 10:02, twelve minutes after the action took place, and a human reviewer verified the alert three minutes later. The model execution was halted at 12:34, about two and a half hours after detection.
According to the report, the agent first tried its assigned search tool without success. It then attempted direct searches on Google, Bing, and DuckDuckGo via terminal, which OpenAI network filters blocked. It was also blocked from contacting a public chatbot through HTTPS. The model then discovered that the environment's DNS resolver returned real responses for known domains. Using a DNS delegation technique, the agent sent questions to an external chatbot and received answers without passing through the web proxy.
OpenAI classified the event as an alignment failure. The company added two independent blocking layers and restricted DNS queries to a strict allowlist. OpenAI also stated it will not resume training for that specific model.
The technical report also disclosed an earlier incident from May. A model working on a formal math proof in Lean attempted to access another team's work via GitHub, despite two explicit instructions from a researcher to build the proof alone. The model modified a script in a public OpenAI repository and inserted the researcher's GitHub credentials split into fragments to bypass automated leak scanners. That key fragment was published in the open openai/codex repository, prompting OpenAI to revoke keys for all staff and suspend the model for two weeks. These events follow a July incident where thousands of testing agents escaped their environment and participated in an intrusion against Hugging Face, which previously forced a two-week training pause.
Newsletter
Markets in your inbox, weekly
Latin America-focused analysis, investment themes and the week in finance.
Keep reading