Benchmark · ReasoningCompetitive

HellaSwag

HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.

Updated 2025-04-15
Models tested
42
Top score
93.7
GPT-4 Turbo
Median
70.1
min 20.9
Top-5 spread
σ 3.7
Competitive

Best score over time · one chart, every benchmark

HELLASWAG8 MODELS · FRONTIER RUNNING MAX0255075100SCORE ↑Sep 24Nov 24Dec 24Feb 25Apr 25RELEASE DATE →benchgecko.ai/benchmark/hellaswag · frontier
Only 8 models have been tested on HellaSwag · not enough history to compute a frontier yet.
Pink dots = frontier records · 0 totalClick to open model page

42 models tested · sorted by score

Same category · related evaluations