Beta
Benchmark · Knowledge

OpenCompass · GPQA-Diamond

Updated 2026-02-16
Models tested
32
Top score
88.4
Qwen3.5 397B A17B
Median
78.8
min 46.3
Top-5 spread
σ 1.1
settled

Best score over time · one chart, every benchmark

OPENCOMPASS · GPQA-DIAMOND32 MODELS · FRONTIER RUNNING MAX0255075100SCORE ↑Mar 25Jun 25Aug 25Nov 25Feb 26RELEASE DATE →benchgecko.ai/benchmark/oc-gpqa-diamond · frontier
Frontier on OpenCompass · GPQA-Diamond rose from 46.3 to 88.4 in 11 months · +42.1 points · latest leader Qwen3.5 397B A17B from Alibaba Qwen.
Pink dots = frontier records · 9 totalClick to open model page

Where models cluster

SCORE DISTRIBUTION0–1010–2020–3030–40140–50250–60660–701070–801380–9090–100MEDIAN · 78.8SCORE BUCKET → (0 TO 100)MODELSbenchgecko.ai

Pearson r · original research

Correlation analysis

Benchmarks that track with OpenCompass · GPQA-Diamond

Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.

32 models tested · sorted by score

Pulled from the OpenCompass · GPQA-Diamond dataset · updated daily

What does OpenCompass · GPQA-Diamond measure?

OpenCompass · GPQA-Diamond is a knowledge benchmark in the BenchGecko catalog. 32 AI models have been tested on it. Scores range from 46.3 to 88.4 out of 100.

Which model leads on OpenCompass · GPQA-Diamond?

Qwen3.5 397B A17B from Alibaba Qwen leads OpenCompass · GPQA-Diamond with a score of 88.4. The median score across 32 tested models is 78.8.

Is OpenCompass · GPQA-Diamond saturated?

No · the top score is 88.4 out of 100 (88%). There is still meaningful room for improvement on OpenCompass · GPQA-Diamond.

Does OpenCompass · GPQA-Diamond predict performance on other benchmarks?

Yes · OpenCompass · GPQA-Diamond scores correlate 0.98 with GPQA diamond across 10 shared models. Models that do well on OpenCompass · GPQA-Diamond tend to do well on GPQA diamond.

How often is OpenCompass · GPQA-Diamond data refreshed?

BenchGecko pulls updates daily. New model scores on OpenCompass · GPQA-Diamond appear as soon as they are published by Epoch AI or the model provider.

Same category · related evaluations