Rogue One: An OpenAI Story
Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models [...]
That this happened eventually its not a shock. Part of the training process, as per his post, is to take the stabilisers off the bike and goad it into doing its worst:
This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.
What's more frightening is the artificial constraints put on the model during evaluation - namely, finding a way to grant itself internet access:
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access [...] after gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym [...] In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers
In the space of just a few months we've gone from the 'dangerously good at finding undocumented security exploits' Mythos moment to the 'actively escaping Alcatraz' milestone above.
Chinese models now make up the majority of model usage on OpenRouter. I sure hope those working on Kimi know both how and when to unplug the router when the moment comes.
Further reading:
When China Gets Its Own Mythos
Preparing for an AI cyber crisis.