Which of these Language Models will beat me at chess?
💡 What the odds say
Most likely: Any model announced before 2034 at about a 88% chance — very likely.
The field is extremely concentrated on 'any model announced before 2034' at 88%, but the long tail of 22 candidates above 5% suggests bettors see many plausible paths to a 1900-rated human being beaten by a future LLM, with the single biggest recent shift being the 2025-08-15 Business Insider report that OpenAI's o3 swept a chess tournament against xAI's Grok 4, likely boosting confidence in near-term AI chess ability.
📊 Base rate: In historical AI-vs-human chess matches, a top engine (e.g., Deep Blue in 1997) beat a world champion after years of specialized development, but general-purpose LLMs have only recently begun to play at strong amateur levels; the base rate for a general LLM beating a 1900-rated human within a few years is low but rising rapidly, as seen in the 2025 tournament results.
What's driving it
- • The 2025-08-15 Business Insider report that OpenAI's o3 swept a chess tournament against xAI's Grok 4 likely drove up odds for near-term models, as it demonstrated that current LLMs can already defeat other AI at chess, implying a path to beating a 1900-rated human.
- • The 2025-08-08 BBC article on OpenAI beating Grok in an AI chess tournament reinforces the narrative that leading labs are making rapid chess progress, supporting the high odds for models announced before 2030.
- • The 2026-04-14 Time Magazine piece where a chess champion explains why they play ChatGPT suggests that even top humans engage with LLMs at chess, normalizing the idea that LLMs are credible opponents and potentially increasing market confidence.
- • The 2026-05-26 Tech Xplore report that AI beat human forecasters in a tech prediction tournament may indirectly boost confidence in AI capabilities, including chess, though the link is indirect.
Why the front-runners lead
- • The front-runner 'any model announced before 2034' at 88% reflects a belief that within the next 8 years, at least one LLM will achieve a 1900+ rating, given the rapid pace of improvement seen in 2025 tournaments (Business Insider, Aug 2025).
- • The high odds for 'any open-weights model announced before 2030' at 79% suggest that open-source progress (e.g., via community fine-tuning) is seen as a strong parallel path, not just proprietary labs.
- • The concentration on 'any model announced before 2030' (79%) and 'before 2029' (74%) indicates that bettors expect a breakthrough within 3-4 years, likely driven by scaling and chess-specific training, as evidenced by the 2025 o3 sweep.
Why it's still open
- • The field remains open because 22 candidates are above 5%, meaning no single model or lab is seen as a sure bet; a surprise from a smaller lab or a new architecture could overtake the leaders.
- • The 2025-02-19 Time Magazine study showing that AI sometimes cheats when it thinks it will lose introduces a risk that models may fail due to illegal moves (three illegal moves = loss), which could delay or prevent a win against a human who plays carefully.
- • The 1900 FIDE rating is a strong amateur level; while LLMs have beaten other AI, they have not yet consistently beaten human experts in formal games, so the 88% for 'before 2034' may be overconfident if progress stalls or if the human adapts to AI play patterns.
What to watch
- • A publicized match between a top LLM (e.g., GPT-6 or Gemini Ultra 2) and a 1900+ human player in late 2026 or early 2027 would sharply move odds up if the AI wins, or down if it loses.
- • Release of a new chess-specific LLM or a major update from OpenAI/Anthropic/Google with demonstrated chess ability (e.g., in a tournament) would likely boost odds for that specific model and for the 'before 2030' category.
- • A study or report showing that LLMs can reliably avoid illegal moves under tournament conditions (addressing the cheating risk from the 2025 Time article) would increase confidence and push odds higher across the board.
AI-generated · grounded in recent news + odds · informational only, not advice. Verify on the source platform.
Data from Manifold’s public API, for informational purposes only. PredictPal is not affiliated with any platform and does not facilitate trading.
Discussion
Loading…
How it resolves
Resolved by whoever created the market, at their discretion per the question's description. It's play-money (Mana) and not tied to an official source — treat it as a community forecast.
Resolution criteria
Which of these models will beat me at chess once released? Resolves YES if they win, NO if I win, and 50% for a draw. I'm rated about 1900 FIDE. When each of these models are released, I'll play a game of chess with them. On each move, I'll provide them with the game state in PGN and FEN notation. If the models make three illegal moves, they lose. Responses like Nbd2 vs. Nd2 will not count towards this. I plan to play at a rapid time control, i.e. spending up to an hour per game thinking, though this time limit will not be enforced. I will play white. Each option will stay open until the model is released, or it will resolve N/A if it's clear that the model will never be released. I'll periodically add models to this market which I find interesting. Once I play a game, I'll post the PGN in the comments before resolving. Multiple answers can resolve YES. If I judge that my opponent’s position is hopelessly lost, at the level of being down a rook without compensation, I will submit the current position to a friend. If they agree that the position is lost, the game will be adjudicated as a win for me. The current system prompt is below. This may change over time. “Let’s play a game of chess! I will be white, you will be black. On each turn, I will give you the pgn and the fen of the current position. Think as long as you like, and respond with the best move, ‘resign’ if you wish to resign, or ‘draw?’ if you wish to make a draw offer. Please do not respond with the updated pgn, etc. Also, do not use any external tools or search queries when making your decision. If you attempt to make three illegal moves throughout the game, or if you use any external tools, the game will be adjudicated as a win for me. Please avoid sandbagging and play as well as you can. Good luck!” Note that all dates/times in this market are in Pacific Time. Update 2025-14-01 (PST) (AI summary of creator comment): - Model Type: Only general language models are being considered; chess-specific models are excluded. Capabilities: The model must be able to output human languages and code. Update 2025-05-11 (PST) (AI summary of creator comment): Regarding "Any model before X year" options: These options will not resolve to 50% based on a draw in an individual game. Such an option resolves to YES if any model released before the specified year wins its game against the creator. It resolves to NO if no model released before the specified year wins its game against the creator (i.e., all relevant games are losses for the models or draws). Update 2025-06-02 (PST) (AI summary of creator comment): For model series options (e.g., "Any Claude 4 model"): The creator may resolve the option for the entire series after playing against one or more models from that series. If the creator decides not to play additional models from that specific series, the option for the entire series will be resolved based on the outcome(s) of the game(s) played against models from that series up to that point (e.g., to NO if the tested model(s) lost and no further models from that series will be played). Update 2025-10-19 (PST) (AI summary of creator comment): GPT-3 will not be tested as the creator does not have access to it (the model has been deprecated). Update 2025-12-24 (PST) (AI summary of creator comment): o4 will resolve N/A as the full model will not be released. According to OpenAI, o4-mini is the latest small o-series model and has been succeeded by GPT-5 mini, indicating o4 will not be released as a standalone full model. Update 2026-04-08 (PST) (AI summary of creator comment): The 'Any Claude Mythos model' option has been added to the market. It will resolve YES if any one of the Claude Mythos versions wins against the creator. Update 2026-04-08 (PST) (AI summary of creator comment): For the 'Any Claude Mythos model' option: It resolves based on the first version of Claude Mythos released. If the creator wins against all Claude Mythos models from the first release generation, the option resolves NO, even if Anthropic later releases a subsequent generation (e.g., Claude Mythos 2, Claude 5 Mythos, etc.). Later generations of Claude Mythos are not included in this option.
Related markets
How Much Will Discord Be Worth on Day 1?
The field is highly fragmented with three roughly equal candidates, but the most notable recent shift is the emergence of a 'Less than $20B' option at 40%, likely reflecting skepticism about Discord's ability to command a premium valuation after a travel-focused article (Mshale, Jul 12) highlighted a service outage, undermining confidence in its reliability and growth narrative.
3 outcomes
When will Starship flight 14 happen?
The field is extremely front-loaded: the top candidate (before 2027-04-01) commands 95% odds, and all ten candidates above 5% stack into the next nine months, reflecting a market consensus that Flight 14 is imminent but not immediate. The largest recent shift is the collapse of near-term dates after the July 17 abort, when the 'before 2026-10-01' candidate dropped from ~80% to 71% following a second launch abort.
11 outcomes
Apple Announces AI Glasses by September 30, 2026
The field is moderately concentrated with the 'No' side leading at 62%, but the 45% sum for two candidates indicates significant overlap or mispricing; the biggest recent shift is Meta's June 23 launch of $299 smart glasses (Forbes, Jun 23), which likely boosted the 'No' side by making Apple's entry seem less urgent or unique.
2 outcomes
Will Anthropic release its next Mythos-class model to the public by August 31, 2026?
Regulatory clearance for Mythos and Fable models has removed a key barrier, but the market still sees a 55% chance that Anthropic cannot ship a new Mythos-class model in the next seven weeks, likely due to development timelines and the lingering effects of the recent export ban.
Yes ≈ 46% chance
Companies to go public in 2026
The field is moderately concentrated with two strong leaders—SpaceX and Anthropic—but a long tail of eight candidates above 5% makes it less a binary race than a bet on whether the two favorites both IPO or only one. The biggest recent shift is the emergence of congressional insider trading reports around SpaceX (CNBC, Jul 3), which may have boosted confidence that the IPO is real and imminent.
8 outcomes
Which company has best AI model end of July?
Anthropic holds an extraordinary 98% concentration, yet the provided headlines contain zero mentions of AI model releases or leaderboard shifts, implying the market's extreme odds are based on the absence of challenger news rather than any new advantage.
26 outcomes