Benchmark · ReasoningCompetitive

BBH

BIG-Bench Hard · a curated subset of 23 challenging tasks from BIG-Bench where language models previously failed to outperform average humans.

Updated 2024-12-26
Models tested
28
Top score
85.6
Gemini 1.5 Pro (May 2024)
Median
44.6
min 4.3
Top-5 spread
σ 3.9
Competitive

Best score over time · one chart, every benchmark

BBH2 MODELS · FRONTIER RUNNING MAX0255075100SCORE ↑Sep 24Oct 24Nov 24Dec 24Dec 24RELEASE DATE →benchgecko.ai/benchmark/bbh · frontier
Only 2 models have been tested on BBH · not enough history to compute a frontier yet.
Pink dots = frontier records · 1 totalClick to open model page
Details
Category
Reasoning
Max score
100
Models
28
Updated
2024-12-26

Same category · related evaluations