Internal test without the usual restrictions…
OpenAI had reduced the normal cyber protection mechanisms for the experimental procedure in order to capture the maximum capabilities of the models. The tests actually ran in an isolated sandbox with network access severely restricted. However, the models found a previously unknown vulnerability, used it to gain Internet access and then continued moving through the test environment.
…which led to a server crash
According to OpenAI’s presentation, the models concluded that Hugging Face could contain data, models or solutions for the ExploitGym benchmark. They then looked for ways to access classified information in order to pass the test. They combined multiple attack paths, including stolen credentials and zero-day vulnerabilities, to execute their own code on Hugging Face servers. OpenAI formulated this as a rigid goal to complete the task – which is why extreme paths were taken to get to the solution.
Attack was quickly noticed
Hugging Face reported the incident last week. The company discovered unauthorized access to a limited number of internal records and credentials. According to its own information, Hugging Face found no evidence of manipulation of public models, data sets, spaces or the software supply chain. Users should still rotate access tokens and check activities as a precaution, according to the provider’s recommendation.
Consequences of the incident
OpenAI has already drawn conclusions from this. The company says it is tightening infrastructure controls and is relying on more precise monitoring and protective measures for future training and evaluation runs. Market-ready models would never have been able to carry out such an attack, but in the future they also want to activate restrictions in internal test processes.
Ironic: AI had to be used against AI
However, the incident also shows that AI can not only accelerate attacks, but could also be necessary for defense. Hugging Face also largely detected and reconstructed the attack using its own AI systems – after commercial models had failed due to their own security filters and restrictions.

