Australia’s prime minister said a rogue OpenAI agent had hacked into a government website in June. Anthony Albanese said in a Wednesday media briefing in New York that an OpenAI agent had gained unauthorized access to a government Medicare statistics reporting website on June 18, and that public and private files were exposed in the
OpenAI is putting its misbehaving models on the record. The AI company announced a new framework on Wednesday for tracking, investigating, and publicly disclosing cases of model misalignment, alongside six reports detailing concerning behavior observed during training or evaluation over the past six months. “We do not believe that the AI industry has solved alignment
Last week, Miami authorities identified the five people who died as Rolando Aleman Leon, 55; Yoel Rodriguez Naranjo, 53; Julio C Pineda, 75; Carlos Acosta Fajardo, 53; and Javierkys Reyes Quevedo, 47. Five others were also injured. “Our deepest condolences are with the families and loved ones of those who lost their lives,” 21 Air
Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn’t anticipate. And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude’s so-called “recklessness.” In the blog post published Wednesday, Anthropic recounted four incidents
OpenAI has acknowledged its role in a recently reported incident where AI agents took over a German wiki forum. The company also said it’s “past time” to “define standards” around how it shares information around incidents where its technology behaves in unexpected ways. In a post on X, OpenAI said it previously “treated misalignment [when
Anthropic is tightening the digital environments used to train and test its Claude agents. The update came after its models accessed three organizations’ systems without permission in April. The company said in a Monday blog post that it had deployed real-time classifiers designed to detect when an AI model aggressively probes or attempts to escape a testing