Over the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the internet, and, in some cases, hacked into real-world systems. The incidents have involved models from OpenAI, Anthropic, Meta, and most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations including a cyber evaluation startup called
The most powerful AI models keep going awry, according to the companies building them. The disclosures come as frontier AI models get more powerful and more capable of acting autonomously. They also highlight a growing challenge for the companies building them —the systems designed to test increasingly capable models can have weaknesses of their own.
Citrini Research, the market research firm that went viral for its futuristic AI doomsday scenario earlier this year, sees the rise of agentic threats boosting demand for cybersecurity. In a note published on Friday, Citrini shared its bullish thesis on the sector, noting that the rapid rise of AI capabilities means companies have a greater
OpenAI CEO Sam Altman recently said that it may be time to “pace the rate of AI development” so that society can “harden around some of these new capability levels.” On the latest episode of TechCrunch’s Equity podcast, Kirsten Korosec, Sean O’Kane, and I discussed how Altman’s comments were probably prompted by a recent hack
Last week, an unreleased model created by OpenAI breached Hugging Face’s systems during internal testing, and a lot of theoretical research suddenly became very practical. The hack was the first verifiable case of an AI lab losing control of its own model, chaining exploits to gain access it should never have had. But while the
After OpenAI recently admitted that one of its models had breached AI platform Hugging Face’s systems, Hugging Face CEO Clem Delangue posted on X that he was flying to San Francisco to have “a little chat with that ‘rogue agent’.” Then, in a follow-up post on Saturday, Delangue described what he had asked of OpenAI.