Benchmark · KnowledgeSettled

LAMBADA

LAMBADA · measures the ability to predict the final word of a passage, requiring broad contextual understanding across long text spans.

Updated 2026-08-13
Models tested
9
Top score
79.8
Falcon-180B
Median
71.8
min 58.4
Top-5 spread
σ 2.8
Competitive

Best score over time · one chart, every benchmark

LAMBADA0 MODELS · FRONTIER RUNNING MAX0255075100SCORE ↑Aug 24Feb 25Aug 25Feb 26Aug 26RELEASE DATE →benchgecko.ai/benchmark/lambada · frontier
Only 0 models have been tested on LAMBADA · not enough history to compute a frontier yet.
Pink dots = frontier records · 0 totalClick to open model page

9 models tested · sorted by score

Same category · related evaluations