JCommonsenseQA
The Frontier
Best score over time · one chart, every benchmark
Full rankings
12 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 95.3 | |
| 2 | 93.7 | |
| 3 | 89.1 | |
| 4 | 87.8 | |
| 5 | 87.7 | |
| 6 | 82.9 | |
| 7 | 78.2 | |
| 8 | 62.4 | |
| 9 | 59.8 | |
| 10 | 52.6 | |
| 11 | 25.5 | |
| 12 | HF SmolLM2 135M Instruct | 17.0 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with JCommonsenseQA
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About JCommonsenseQA
What does JCommonsenseQA measure?
JCommonsenseQA is a knowledge benchmark in the BenchGecko catalog. 12 AI models have been tested on it. Scores range from 17.0 to 95.3 out of 100.
Which model leads on JCommonsenseQA?
DeepSeek R1 Distill Qwen 32B from DeepSeek leads JCommonsenseQA with a score of 95.3. The median score across 12 tested models is 80.6.
Is JCommonsenseQA saturated?
Yes · the top model on JCommonsenseQA has reached 95.3 out of 100, within 5% of the theoretical ceiling. This benchmark is approaching saturation and may be replaced by a harder successor.
Does JCommonsenseQA predict performance on other benchmarks?
Yes · JCommonsenseQA scores correlate 0.91 with LLM-JP · Overall across 12 shared models. Models that do well on JCommonsenseQA tend to do well on LLM-JP · Overall.
How often is JCommonsenseQA data refreshed?
BenchGecko pulls updates daily. New model scores on JCommonsenseQA appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 12
- Updated
- 2025-01-20
Top on JCommonsenseQA
DeepSeek R1 Distill Qwen 32B · 95.3DeepSeek R1 Distill Qwen 14B · 93.7Qwen2 7B Instruct · 89.1Qwen2 VL 7B Instruct · 87.8Meta Llama 3 8B Instruct · 87.7More knowledge benchmarks
Same category · related evaluations