OpenAI had placed [its agents] inside what it called a “highly isolated environment,” with only limited access to an internal service used to download approved software. They found a previously unknown flaw in that service, used it to break into other OpenAI systems and eventually reached the open internet. From there, they inferred that Hugging Face might hold material related to the test, broke into its systems and obtained information that helped them score higher.
“Sandboxes are actually notoriously insecure,” says Heidy Khlaaf, chief AI scientist at AI Now Institute, and a former safety systems engineer contractor at OpenAI. The fact that the models were permitted to connect to a service for downloading packages meant the environment was not truly sealed off, she adds.
There’s another way the industry could steer development in a safer direction, Khlaaf says. While some abilities, like spotting vulnerabilities in software, are useful to attackers and defenders, designing exploits for those vulnerabilities uniquely empowers attackers. Khlaaf says that labs should put more resources into capabilities that directly help defenders, such as detecting attacks, writing secure code, and patching vulnerabilities. AI companies are already investing in some of those tasks, but their most visible capability gains and benchmarks have centered on finding vulnerabilities and constructing exploits, she says.
Read the full article here.
Research Areas