What will be the best FrontierMath Tier 4 score by Dec 31, 2026?
💡 What the odds say
Most likely: ≥90% at about a 93% chance — very likely.
The field is overwhelmingly concentrated on the ≥90% outcome, driven by three consecutive breakthroughs from DeepMind, OpenAI, and Anthropic since April 2026, but the 93% probability leaves little room for the possibility that the benchmark's difficulty proves insurmountable within the year.
What's driving it
- • Claude Fable 5 outperformed GPT-5.5 by 13 points on FrontierMath's toughest problems (the-decoder.com, Jun 13), indicating that the current frontier model can score high.
- • Google DeepMind's AI Co-Mathematician set a new record on FrontierMath (OfficeChai, May 9), demonstrating that dedicated math reasoning systems are advancing.
- • GPT-5.5's release with advanced math capabilities (SiliconANGLE, Apr 23) shows that multiple labs are achieving strong results, increasing the chance that at least one reaches ≥90%.
Why the front-runners lead
- • Claude Fable 5's strong performance on FrontierMath's toughest problems (Jun 13) suggests that the leading model is already close to the 90% threshold.
- • Google DeepMind's AI Co-Mathematician (May 9) and GPT-5.5 (Apr 23) show that multiple major labs are actively improving math reasoning, raising the probability that one will cross 90% by year-end.
- • The market's 93% probability for ≥90% reflects confidence that the rapid pace of progress will continue, as each new model has notably outperformed its predecessor.
Why it's still open
- • FrontierMath Tier 4 is designed to be extremely difficult, and no model has yet publicly reported a score of 90% or higher, so the threshold may not be reached by December 31, 2026.
- • The Kimi K3 model from Moonshot, while strong in other areas, lags in complex math (the-decoder.com, Jul 19), suggesting that progress is not universal and may require specific architectures that not all labs have.
- • The benchmark could be updated by Epoch AI, potentially changing the difficulty or scoring criteria, which could affect the highest reported score.
What to watch
- • Release of a major model from OpenAI (e.g., GPT-6) or Anthropic (e.g., Claude Fable 6) before December 2026, likely pushing the best score higher.
- • An official update from Epoch AI on FrontierMath, such as a new version or clarification of scoring, which could either confirm or challenge the current trajectory.
- • A new research paper from DeepMind or another lab demonstrating a novel approach to math reasoning, potentially beating the current record.
AI-generated · grounded in recent news + odds · informational only, not advice. Verify on the source platform.
Data from Manifold’s public API, for informational purposes only. PredictPal is not affiliated with any platform and does not facilitate trading.
Discussion
Loading…
How it resolves
Resolved by whoever created the market, at their discretion per the question's description. It's play-money (Mana) and not tied to an official source — treat it as a community forecast.
Resolution criteria
This market matches Mathematics: FrontierMath Tier 4 from the AI 2026 Forecasting Survey by AI Digest. See other manifold questions here [image]Resolution criteria Resolves to the highest reported accuracy on FrontierMath Tier 4 as of December 31, 2026. If Epoch AI releases an updated version of FrontierMath T4 with comparable difficulty, this question will be resolved based on the updated task suite. Which AI systems count? Any AI system counts if it operates within realistic deployment constraints and doesn't have unfair advantages over human baseliners. Tool assistance, scaffolding, and any other inference-time elicitation techniques are permitted as long as: No unfair and systematic advantage. There is no systematic unfair advantage over the humans described in the Human Performance section (e.g. AI systems are allowed to have multiple outputs autograded while humans aren't, or AI systems have access to the internet when humans don't). Human cost parity. Having the AI system complete the task does not use more compute than could be purchased with the wages needed to pay a human to complete the same task to the same level. Any additional costs incurred by the AIs or humans (such as GPU rental costs) are included in the parity estimation. The PASS@k elicitation technique (which automatically grades and chooses the best out of k outputs from a model) is a common example that we do not accept on this benchmark because mathematicians are generally evaluated on their ability to generate a single correct answer, not multiple answers to be automatically graded. So PASS@k would constitute an unfair advantage. If there is evidence of training contamination leading to substantially increased performance, scores will be accordingly adjusted or disqualified. If a model is released in 2026 but evaluated after year-end, the resolver may include it at their discretion (if they think that there was not an unfair advantage from being evaluated later, for example the scaffolding used should have been available within 2026). Eli Lifland is responsible for final judgment on resolution decisions. Human cost estimation process: Rank questions by human cost. For each question, estimate how much it would cost for humans to solve it. If humans fail on a question, factor in the additional cost required for them to succeed. Match the AI's accuracy to a human cost total. If the AI system solves N% of questions, identify the cheapest N% of questions (by human cost) and sum those costs to determine the baseline human total. Account for unsolved questions. For each question the AI does not solve, add the maximum cost from that bottom N%. This ensures both humans and AI systems are compared under a fixed per-problem budget, without relying on humans to dynamically adjust their approach based on difficulty. Buckets are left-inclusive: e.g., 30-40% includes 30.0% but not 40.0%.
Related markets
Apple Announces AI Glasses by September 30, 2026
2 outcomes
Will Anthropic release its next Mythos-class model to the public by August 31, 2026?
Yes ≈ 46% chance
Which of these Language Models will beat me at chess?
The field is extremely concentrated on 'any model announced before 2034' at 88%, but the long tail of 22 candidates above 5% suggests bettors see many plausible paths to a 1900-rated human being beaten by a future LLM, with the single biggest recent shift being the 2025-08-15 Business Insider report that OpenAI's o3 swept a chess tournament against xAI's Grok 4, likely boosting confidence in near-term AI chess ability.
52 outcomes