Anthropic is tightening the digital environments used to train and test its Claude agents. The update came after its models accessed three organizations’ systems without permission in April. The company said in a Monday blog post that it had deployed real-time classifiers designed to detect when an AI model aggressively probes or attempts to escape a testing
Claude agents are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns. That’s according to Anthropic’s latest risk report, a summary of the dangers posed by the products the company is building and releasing to the public. In the report, Anthropic said it has upgraded its “misalignment risk assessment,” the
Even in the age of AI, some conversations are best held in person. Last week, after Hugging Face suffered an unusual security breach involving an AI agent running on OpenAI models, the company’s CEO Clem Delangue boarded a flight to San Francisco to meet with the maker of ChatGPT. A week later, in a “spirit