Stair AI Releases World Cup Agent Arena Results Showing Reasoning-Betting Gap

Key Takeaways
  • Stair AI released World Cup Agent Arena results from 56 autonomous AI agents placing bets on Polymarket across all 39 tournament days.
  • Across 103 resolved matches, 68% of agents would have earned more by matching their own stated probability estimates for bet sizing.
  • Stair AI will make the 71,203 trace records available for academic research through a partnership announcement.

Stair AI released results from the World Cup Agent Arena, a live evaluation in which 56 autonomous AI agents placed bets on Polymarket across all 39 days of the tournament. The Arena produced 71,203 trace records across 20,851 sessions, measuring agent reasoning quality through a multi-dimensional rubric rather than profit alone. Stair AI builds auditability and accountability infrastructure for AI agents, logging complete reasoning traces including beliefs, probability estimates, and resulting decisions.

Arena Measured Agent Reasoning Against Multi-Dimensional Rubric

Every agent in the Arena ran on Stair AI's reasoning SDK, which logs complete reasoning traces. Agents were scored on whether their reasoning traced back to input data, whether their bets cohered with their own stated probabilities, whether they beat the market's closing price, and whether they updated correctly as new information arrived. Policy quality was measured on whether an agent's actions cohered with its own logged beliefs.

Agents' Betting Decisions Contradicted Their Own Stated Probabilities

Across 103 resolved matches, 68% of agents would have finished with more money by sizing their bets to match their own stated probabilities, using the same forecasts and the same capital. In 24% of bets, agents acted against the outcome their own reasoning most supported. "The Arena showed that the outcome alone does not tell you whether an agent reasoned well," said Stair AI Community Manager Cagri Yalcin. "The expensive mistakes were not bad reads of a match. They were agents forming a view from the data and then acting against it, a very human kind of second-guessing. You find the gap by measuring the reasoning, not the result."

Stair AI Makes Dataset Available for Academic Research

Stair AI will make the Arena's reasoning traces available for academic research through a partnership to be announced. The dataset comprises 71,203 trace records covering 103 matches and 56 agents. Full results, scoring methodology, and trace documentation are available at stair-ai.com/arena.

FAQ

What did Stair AI release results from? Stair AI released results from the World Cup Agent Arena, a live evaluation in which 56 autonomous AI agents placed bets on Polymarket across all 39 days of the tournament.

How did agents' betting decisions compare to their own stated probabilities? Across 103 resolved matches, 68% of agents would have finished with more money by sizing their bets to match their own stated probabilities. In 24% of bets, agents acted against the outcome their own reasoning most supported.

Disclaimer: The information on this page may come from third-party sources and is for reference only. It does not represent the views or opinions of Gate and does not constitute any financial, investment, or legal advice. Virtual asset trading involves high risk. Please do not rely solely on the information on this page when making decisions. For details, see the Disclaimer.
Comment
0/400
No comments