What’s the least impressive thing you’re very sure AI still won’t be able to do before August 2027? [read description]
💡 What the odds say
Most likely: Generating labeled diagrams of some arbitrary device(s) (within reason) at about a 100% chance — almost certain.
The field is wide open with 39 options and the top three barely above 5%, but generating labeled diagrams is at 100% because it is the most concrete and verifiable candidate, while the next tier clusters around 75-88%, reflecting uncertainty about subjective criteria like natural conversations and taxes.
What's driving it
- • Generating labeled diagrams surged to 100% likely due to its objective verifiability and the market's rule that options must be realistically verifiable by the moderator (see resolution details), making it a safe bet.
- • No clear catalyst explains the 88% odds for natural conversations and taxes; the lack of recent headlines about AI conversation benchmarks or tax automation suggests traders are anchoring on known AI limitations from earlier reports.
- • Beating a Pokémon game glitchless at 100% may be driven by the recent proven success of ClaudePlaysPokémon, which showed AI can play games with assistance, making the glitchless constraint a plausible near-term achievement.
- • The 82% for turning $1k into $1.2k in a year reflects no direct headline catalyst, but ties into the general trend of AI trading and financial manipulation being hard to verify, as seen in the least reliable EVs article (BGR, Feb 7) which highlights trust issues in tech claims.
Why the front-runners lead
- • Generating labeled diagrams leads because it is a well-defined, easy-to-verify task that matches the market's emphasis on realistic verification by the moderator.
- • Natural conversations at 88% may be high because users assume current AI like Claude can already mimic natural flow, though the 'human knows it's an AI' condition is a subjective bar that could fail.
- • Booking airline tickets at 85% reflects the belief that simple instructions map to existing automated booking systems, but real-world variability in prices and policies remains a hurdle.
Why it's still open
- • Recognizing sarcasm at 78% is vulnerable because it requires nuanced human-like reasoning beyond current NLP benchmarks, and no recent headline shows breakthrough in sarcasm detection.
- • Solving novel cryptic crossword clues at 79% could be overturned by the difficulty of generalizing AI reasoning to novel wordplay, as crossword experts often rely on cultural and contextual knowledge.
- • The field is open because 36 of 39 options are above 5%, meaning many bettors think mainstream tasks like taxes or sarcasm will fail, which is plausible given the subjective resolution criteria and lack of concrete progress in recent headlines.
What to watch
- • Any major AI model release (e.g., GPT-5, Claude 4) before market close could shift odds up for natural conversations and taxes if benchmarks show improvement, but down for sarcasm and crossword if limitations persist.
- • A public verification of an AI completing end-to-end taxes by a reputable source would push that option above 90%, while a failed experiment would drop it.
- • The August 2027 deadline and the requirement for realistic verification by the moderator mean that concrete demonstrations (e.g., a streamed Pokémon game win or a published tax filing) are catalysts, likely increasing odds for front-runners as deadlines approach.
AI-generated · grounded in recent news + odds · informational only, not advice. Verify on the source platform.
Data from Manifold’s public API, for informational purposes only. PredictPal is not affiliated with any platform and does not facilitate trading.
Discussion
Loading…
How it resolves
Resolved by whoever created the market, at their discretion per the question's description. It's play-money (Mana) and not tied to an official source — treat it as a community forecast.
Resolution criteria
OPTIONS RESOLVE YES IF THEY HAPPEN 1 option per person, but if you can credit the prediction about this question to a public person you can add it. Interpret the question the way you find most reasonable. You can explain your choice in the comments! IMPORTANT: the prediction must be realistically verifiable by me (can involve some searching or simple experiment), @Bayesian, in the event that you are not reachable at time of market close. I will N/A options where this is not the case. If abs(your mana net worth) < 5000, I’ll cover the cost of your option if you ping me or DM me. Some details: The spirit of the market is that if an option / benchmark / stated prediction is achieved via methods that would be deemed scientific malpractice, obvious trickery, or deception, it will not count as a valid resolution. For example, "AI 10x's a portfolio in a year" would not count if 10 different instances of the AI try the same challenge with their own pot of money and only one of them succeeds, and the other 9 go to 0. if the task is simple for specialized AI systems to solve today, we can safely assume the intent is to only count chatbot-style systems Inspired by @liron tweet [tweet]Update 2025-07-21 (PST) (AI summary of creator comment): The creator has specified how different types of AI will be considered: The market applies to any AI system, not exclusively to LLMs. However, for options that implicitly refer to a specific AI capability (e.g., 'jailbreaking' a chatbot), the market will be judged based on the most competent systems of that relevant type. Update 2025-07-21 (PST) (AI summary of creator comment): The creator has clarified how options will be judged based on their phrasing: If an option describes a capability, it will be resolved based on whether an AI has that capability, provided it is safe and practical to test. If an option describes an action, it will be resolved based on whether an AI actually performs that action. Update 2025-07-21 (PST) (AI summary of creator comment): The creator has specified their process for determining if an AI has a certain capability: The creator will attempt to personally elicit the behavior from a relevant AI system and will also search for public online evidence. If evidence of the capability is not found through these methods, it will be concluded that the AI cannot do the action. Update 2025-07-21 (PST) (AI summary of creator comment): The creator has clarified that the required frequency of an action depends on the context of the option: For one-off events, a single occurrence is sufficient for the option to resolve YES. For tasks that imply a skill (e.g. mathematical calculations), a single success by random chance is not sufficient. These will require some level of consistency to be demonstrated. Update 2025-07-21 (PST) (AI summary of creator comment): The creator has confirmed that an option is considered acceptable and verifiable even if it includes a negative constraint on the AI's method, such as requiring a task to be performed without using tools (e.g., without writing and executing code). Update 2025-07-22 (PST) (AI summary of creator comment): The creator has provided an example of how they will interpret options that are ambiguously phrased about the type of AI. If an option is broad enough to include any AI, the creator may test it against very simple systems. For example, for an option involving 'learning', a simple database AI memorizing information could be considered sufficient to meet the criteria, causing the option to resolve YES. Update 2025-07-22 (PST) (AI summary of creator comment): The creator has stated that the distinction between an AI's capability in text versus in speech is an important one that will be considered during resolution. Update 2025-07-23 (PST) (AI summary of creator comment): The creator has specified that the duration of the verification process is a factor in whether an option is considered realistically testable. Options that require a long period to verify (e.g., one year) are considered unverifiable and will be resolved to N/A. Update 2025-07-23 (PST) (AI summary of creator comment): In a discussion about an answer involving an AI outperforming human forecasters, the creator has clarified their approach to such ambiguous claims: Phrasings like "better than human experts" are considered hard to verify due to ambiguity (e.g., better than the worst, average, or best expert?). A more concrete and verifiable benchmark would be required for resolution. The creator suggested a possibility could be comparing forecasting bots against human averages on a platform like Metaculus. Update 2025-07-23 (PST) (AI summary of creator comment): The creator has resolved a specific answer to N/A, stating that its meaning "changed too much" during a discussion in the comments. This indicates that other answers may be resolved to N/A if their definition is significantly altered after being submitted. Update 2025-07-23 (PST) (AI summary of creator comment): In a discussion about an answer involving an AI performing a task in “some area”, the creator has clarified their interpretation: The condition may be considered met if the AI can perform the task in any specific area, even a simple or “economically useless niche”. Update 2025-07-23 (PST) (AI summary of creator comment): In a discussion about an answer that is difficult for the creator to personally test (e.g., an AI making a large amount of money over a year), the creator has proposed an alternative to resolving it to N/A: The resolution can be based on the existence of a credible public report about the event by the market close date. If no such report is found, the event will be considered to not have happened (i.e., the answer will resolve NO). Update 2025-07-24 (PST) (AI summary of creator comment): In a discussion about an answer related to an AI recognizing sarcasm, the creator clarified that answers may be considered too ambiguous for verification if they do not specify the modality to be tested (e.g., text, voice, or both). Update 2025-07-24 (PST) (AI summary of creator comment): In response to a question about how regulatory limitations will be judged, the creator has clarified: Options that are limited by regulation are acceptable. To make the resolution dependent on an AI's legal status to perform a task, the option should be phrased explicitly, for example, "can legally do X". Otherwise, the option will be judged based on the AI's technical capability to perform the task, ignoring regulatory constraints. Update 2025-07-25 (PST) (AI summary of creator comment): In a discussion about an answer involving financial returns, the creator clarified how they will assess the validity of a success: A single attempt by a single entity (e.g., a lab) that succeeds will generally be counted for resolution, unless the creator deems it suspicious. This is in contrast to the existing rule where an outcome achieved by only one of many AI instances attempting the same challenge will not be counted. Update 2025-07-25 (PST) (AI summary of creator comment): In a discussion about an answer involving an AI making a financial return, the creator has clarified their interpretation of specific terms: An action is not considered "independent" if most of the task is set up for the AI (e.g., being given a pre-stocked vending machine to run). For financial returns, the resolution will be based on the absolute return achieved. For example, an AI making a 20% return is a success, even if a market index like the S&P 500 grew by more in the same period. Update 2025-07-25 (PST) (AI summary of creator comment): The creator has clarified the process for defining verification criteria for individual answers: The creator of an answer can provide input on the verification procedure for their own submission. The market creator may agree to add specific verification requirements (e.g., that results must be from a peer-reviewed study) to an individual answer if proposed by that answer's creator. The market creator remains the final arbitrator on all resolutions. Update 2025-07-30 (PST) (AI summary of creator comment): In a discussion about an ambiguous answer, the creator has stated their intention to resolve it to N/A. However, they will first invite the answer's submitter to provide a more detailed and verifiable version to avoid this resolution. Update 2025-11-03 (PST) (AI summary of creator comment): For the answer "teleoperate a robot to tidy up random kitchens - Gary Marcus": Resolution will be based on a "you know it when you see it" standard Human guidance through the kitchen is acceptable (the AI does not need to one-shot infer where everything goes)
Related markets
How Much Will Discord Be Worth on Day 1?
The field is highly fragmented with three roughly equal candidates, but the most notable recent shift is the emergence of a 'Less than $20B' option at 40%, likely reflecting skepticism about Discord's ability to command a premium valuation after a travel-focused article (Mshale, Jul 12) highlighted a service outage, undermining confidence in its reliability and growth narrative.
3 outcomes
When will Starship flight 14 happen?
The field is extremely front-loaded: the top candidate (before 2027-04-01) commands 95% odds, and all ten candidates above 5% stack into the next nine months, reflecting a market consensus that Flight 14 is imminent but not immediate. The largest recent shift is the collapse of near-term dates after the July 17 abort, when the 'before 2026-10-01' candidate dropped from ~80% to 71% following a second launch abort.
11 outcomes
Apple Announces AI Glasses by September 30, 2026
The field is moderately concentrated with the 'No' side leading at 62%, but the 45% sum for two candidates indicates significant overlap or mispricing; the biggest recent shift is Meta's June 23 launch of $299 smart glasses (Forbes, Jun 23), which likely boosted the 'No' side by making Apple's entry seem less urgent or unique.
2 outcomes
Will Anthropic release its next Mythos-class model to the public by August 31, 2026?
Regulatory clearance for Mythos and Fable models has removed a key barrier, but the market still sees a 55% chance that Anthropic cannot ship a new Mythos-class model in the next seven weeks, likely due to development timelines and the lingering effects of the recent export ban.
Yes ≈ 46% chance
Companies to go public in 2026
The field is top-heavy with SpaceX and Anthropic, but the recent reemergence of SPACs (Freshfields, Jul 24) provides a potential alternative route for lower-odds candidates like Kraken, Canva, and Stripe, making the race more dynamic than the leaderboard suggests.
8 outcomes
Which of these Language Models will beat me at chess?
The field is extremely concentrated on 'any model announced before 2034' at 88%, but the long tail of 22 candidates above 5% suggests bettors see many plausible paths to a 1900-rated human being beaten by a future LLM, with the single biggest recent shift being the 2025-08-15 Business Insider report that OpenAI's o3 swept a chess tournament against xAI's Grok 4, likely boosting confidence in near-term AI chess ability.
52 outcomes