OpenBookQA
OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
26 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 84.0 | |
| 2 | 84.0 | |
| 3 | 83.2 | |
| 4 | 81.3 | |
| 5 | 81.1 | |
| 6 | 81.1 | |
| 7 | 76.8 | |
| 8 | 73.1 | |
| 9 | 71.5 | |
| 10 | 64.8 | |
| 11 | 52.3 | |
| 12 | U PaLM 2-M | 43.2 |
| 13 | 42.7 | |
| 14 | 41.9 | |
| 15 | U PaLM 2-S | 41.6 |
| 16 | U MPT-30B | 36.0 |
| 17 | 32.3 | |
| 18 | 30.1 | |
| 19 | U XGen-7B | 20.3 |
| 20 | U RedPajama-INCITE-7B-Base | 20.0 |
| 21 | U Dolly 2.0-12b | 18.9 |
| 22 | 18.7 | |
| 23 | 16.3 | |
| 24 | 14.4 | |
| 25 | U vicuna-13b-v1.1 | 10.7 |
| 26 | U stablelm-tuned-alpha-7b | 9.9 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with OpenBookQA
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About OpenBookQA
What does OpenBookQA measure?
OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting. 26 AI models have been tested on it. Scores range from 9.9 to 84.0 out of 100.
Which model leads on OpenBookQA?
phi-3-mini 3.8B from Microsoft leads OpenBookQA with a score of 84.0. The median score across 26 tested models is 42.3.
Is OpenBookQA saturated?
No · the top score is 84.0 out of 100 (84%). There is still meaningful room for improvement on OpenBookQA.
Does OpenBookQA predict performance on other benchmarks?
Yes · OpenBookQA scores correlate 0.87 with ANLI across 10 shared models. Models that do well on OpenBookQA tend to do well on ANLI.
How often is OpenBookQA data refreshed?
BenchGecko pulls updates daily. New model scores on OpenBookQA appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 26
- Updated
- 2024-07-16
Top on OpenBookQA
phi-3-mini 3.8B · 84.0phi-3-small 7.4B · 84.0phi-3-medium 14B · 83.2GPT-3.5 Turbo (older v0613) · 81.3Mixtral 8x7B · 81.1More knowledge benchmarks
Same category · related evaluations