For a year now, AI security testing company Andon Labs has tasked frontier models with various real-world tasks to determine how well they perform as agents operating for long periods without human supervision. On Wednesday, Andon published a new installment about how things are going in its Vending-Bench research, where the lab has cutting-edge models