When will AI pass @jim's "agents benchmark"?
💡 What the odds say
Most likely: 2026 at about a 78% chance — likely.
Data from Manifold’s public API, for informational purposes only. PredictPal is not affiliated with any platform and does not facilitate trading.
Discussion
Loading…
How it resolves
Resolved by whoever created the market, at their discretion per the question's description. It's play-money (Mana) and not tied to an official source — treat it as a community forecast.
Resolution criteria
Resolves to the year during which the "agents benchmark" is first solved. The benchmark involves an AI being given the ASCII art shown below and being asked to colour each of the depicted figures in a different colour. If an AI succeeds at least half the time it is considered to have passed the benchmark. The output should be HTML or HTML and CSS. The prompt must consist of no more than two English language sentences (along with the ASCII art itself). The AI must have a pass rate of at least 50% for the solution to qualify. [image] o o__ __o o__ __o__/_ o o ____o__ __o____ o__ __o <|> /v v\ <| v <|\ <|> / \ / \ /v v\ / \ /> <\ < > / \\o / \ \o/ /> <\ o/ \o o/ | \o/ v\ \o/ | _\o____ <|__ __|> <| _\__o__ o__/_ | <\ | < > \_\__o__ / \ \\ | | / \ \o / \ | \ o/ \o \ / <o> \o/ v\ \o/ o \ / /v v\ o o | | <\ | <| o o /> <\ <\__ __/> / \ _\o__/_ / \ < \ / \ <\__ __/> Update 2025-02-16 (PST) (AI summary of creator comment): 333 Characters Limit Update The prompt (excluding the ASCII art) must contain no more than 333 characters in total. It must still consist of no more than two English language sentences.
ⓘ A market settles under its own written rules, which can lag what looks decided in the news — so the price may not move to 100% the moment an outcome seems obvious.
Related markets
Apple Announces AI Glasses by September 30, 2026
2 outcomes
Will Anthropic release its next Mythos-class model to the public by August 31, 2026?
Yes ≈ 46% chance
Which of these Language Models will beat me at chess?
The field is extremely concentrated on 'any model announced before 2034' at 88%, but the long tail of 22 candidates above 5% suggests bettors see many plausible paths to a 1900-rated human being beaten by a future LLM, with the single biggest recent shift being the 2025-08-15 Business Insider report that OpenAI's o3 swept a chess tournament against xAI's Grok 4, likely boosting confidence in near-term AI chess ability.
52 outcomes