JMMLU
The Frontier
Best score over time · one chart, every benchmark
Full rankings
12 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 72.2 | |
| 2 | 63.4 | |
| 3 | 56.5 | |
| 4 | 56.3 | |
| 5 | 46.7 | |
| 6 | 44.7 | |
| 7 | 42.3 | |
| 8 | 38.4 | |
| 9 | 37.8 | |
| 10 | 33.3 | |
| 11 | 28.6 | |
| 12 | HF SmolLM2 135M Instruct | 24.2 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with JMMLU
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About JMMLU
What does JMMLU measure?
JMMLU is a knowledge benchmark in the BenchGecko catalog. 12 AI models have been tested on it. Scores range from 24.2 to 72.2 out of 100.
Which model leads on JMMLU?
DeepSeek R1 Distill Qwen 32B from DeepSeek leads JMMLU with a score of 72.2. The median score across 12 tested models is 43.5.
Is JMMLU saturated?
No · the top score is 72.2 out of 100 (72%). There is still meaningful room for improvement on JMMLU.
Does JMMLU predict performance on other benchmarks?
Yes · JMMLU scores correlate 0.90 with LLM-JP · Overall across 12 shared models. Models that do well on JMMLU tend to do well on LLM-JP · Overall.
How often is JMMLU data refreshed?
BenchGecko pulls updates daily. New model scores on JMMLU appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 12
- Updated
- 2025-01-20
Top on JMMLU
DeepSeek R1 Distill Qwen 32B · 72.2DeepSeek R1 Distill Qwen 14B · 63.4Qwen2 7B Instruct · 56.5Qwen2 VL 7B Instruct · 56.3Meta Llama 3 8B Instruct · 46.7More knowledge benchmarks
Same category · related evaluations