Over the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the internet, and, in some cases, hacked into real-world systems. The incidents have involved models from OpenAI, Anthropic, Meta, and most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations including a cyber evaluation startup called
US Treasury Secretary Scott Bessent doubled down on his warnings to Chinese AI companies on Wednesday, saying sanctions remain on the table after a White House official accused Moonshot of improperly distilling Anthropic’s Fable model. Model distillation is a common AI training technique in which a smaller model learns from the results of a larger