How do you measure whether an AI is really good at predicting the future? The CTO of startup Obside recently emailed me with a fascinating real-world benchmark. Instead of giving the AI models another standardized test, Obside had ChatGPT, Gemini, Claude, Grok, Mistral, DeepSeek, and Kimi bet on World Cup matches using live odds from
How do you measure whether an AI is really good at predicting the future? The CTO of startup Obside recently emailed me with a fascinating real-world benchmark.
Instead of giving the AI models another standardized test, Obside had ChatGPT, Gemini, Claude, Grok, Mistral, DeepSeek, and Kimi bet on World Cup matches using live odds from Polymarket.
An hour before kickoff, each model goes into agent mode, researches teams, injuries and other public information, then decides how much of their $10,000 virtual budget to bet.
When I checked after the semi-finals on Thursday, French open source favorite Mistral was leading the field, followed by OpenAI’s GPT 5.5 and DeepSeek’s V4. Claude Opus 4.8, meanwhile, sat firmly at the bottom, the only model in red. Maybe Anthropic’s AI is simply too ethical to be a gamer?
Check the current classification here.
This fun exercise measures something many AI benchmarks can’t: judging under conditions of uncertainty. That’s a topic I’ve explored before. Last year, I wrote about ChatGPT participating in a secret forecasting tournament run by economists, where it did no better than the average human contestant.
Betting on football isn’t the same, but it’s another clever way to test whether AI can turn online information into profitable predictions when no one knows the answer yet.
Subscribe to BI’s Tech Memo newsletter here. Contact me by email at abarr@businessinsider.com.
