WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
84 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with WeirdML
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About WeirdML
What does WeirdML measure?
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. 84 AI models have been tested on it. Scores range from 1.7 to 87.8 out of 100.
Which model leads on WeirdML?
Claude Fable 5 from Anthropic leads WeirdML with a score of 87.8. The median score across 84 tested models is 42.1.
Is WeirdML saturated?
No · the top score is 87.8 out of 100 (88%). There is still meaningful room for improvement on WeirdML.
Does WeirdML predict performance on other benchmarks?
Yes · WeirdML scores correlate 0.92 with GPQA diamond across 73 shared models. Models that do well on WeirdML tend to do well on GPQA diamond.
How often is WeirdML data refreshed?
BenchGecko pulls updates daily. New model scores on WeirdML appear as soon as they are published by Epoch AI or the model provider.
- Category
- Code
- Max score
- 100
- Models
- 84
- Updated
- 2026-06-09
Top on WeirdML
Claude Fable 5 · 87.8Claude Opus 4.8 · 82.9GPT-5.3-Codex · 79.3Claude Opus 4.6 · 78.0GPT-5.4 · 77.7More code benchmarks
Same category · related evaluations