Anthropic says its new model already found thousands of “high-severity” vulnerabilities across the tech landscape at a level that surpasses human experts.

[…]

But there are significant doubts about those claims, and Heidy Khlaaf, chief AI scientist at the AI Now Institute, wasn’t impressed. She’s spent her career building and auditing the exact kinds of code analysis tools that Anthropic suggests it’s surpassed. She’s also worked on digital safety in nuclear facilities.

Khlaaf says the biggest red flag was the lack of false positive rates – an industry-standard measure of how often a security tool flags something that isn’t a real problem. “This is not some unknown metric,” Khlaaf says. “This is kind of the largest indicator of how useful your tool is.” Anthropic didn’t mention it and sidestepped the question when I asked for comment. Nor did Anthropic measure Mythos against existing tools that security engineers have relied on for decades.

None of this is to say that the threat is imaginary. “Mythos might be capable,” Khlaaf says. AI tools are genuinely well-suited for scanning massive code bases, and automatically finding security vulnerabilities is a real and pressing danger. But Khlaaf is skeptical about Anthropic’s claims without being able to substantiate them. “I think there are a lot of cracks in this narrative that Mythos is all powerful, we can’t release it.”

Read the article here.

Research Areas