Recent model releases have driven strong trader consensus on MathArena benchmarks, with Anthropic's Claude Opus 5 (max) posting the highest verified score of 84.4% on recent competition problems in July 2026, ahead of OpenAI's GPT-5.6-Sol at 79.7%. Frontier labs continue advancing reasoning through extended chain-of-thought techniques and specialized training on olympiad-level datasets, narrowing gaps with human experts on AIME and Putnam-style tasks. Competitive dynamics favor closed models from Anthropic and OpenAI, though open-weight entries like Moonshot's Kimi K3 trail at around 70%. With four months remaining until year-end, upcoming releases or fine-tunes from Google, xAI, or Meta could push scores higher if they demonstrate gains on uncontaminated ArXivMath or live olympiad evaluations. Traders monitor official announcements for capability thresholds that would resolve the market.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$114,845 Vol.
1575
73%
1600
28%
$114,845 Vol.
1575
73%
1600
28%
Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Market Opened: Apr 2, 2026, 6:07 PM ET
Resolver
0x65070BE91...Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Resolver
0x65070BE91...Recent model releases have driven strong trader consensus on MathArena benchmarks, with Anthropic's Claude Opus 5 (max) posting the highest verified score of 84.4% on recent competition problems in July 2026, ahead of OpenAI's GPT-5.6-Sol at 79.7%. Frontier labs continue advancing reasoning through extended chain-of-thought techniques and specialized training on olympiad-level datasets, narrowing gaps with human experts on AIME and Putnam-style tasks. Competitive dynamics favor closed models from Anthropic and OpenAI, though open-weight entries like Moonshot's Kimi K3 trail at around 70%. With four months remaining until year-end, upcoming releases or fine-tunes from Google, xAI, or Meta could push scores higher if they demonstrate gains on uncontaminated ArXivMath or live olympiad evaluations. Traders monitor official announcements for capability thresholds that would resolve the market.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions