Gecko Tests · Knowledge HorizonLatest run 2026-10-04 · 7 models · 33 months · 264 questions

Knowledge Horizon · Where Does the Model’s World End?

Labs publish a training cutoff. We measure it: questions about each month's world events since January 2024, answered with no web access. The horizon is the last month the model still knows about.

ModelMeasured horizonClaimed cutoffReleasedAccuracy by month2024 accuracy
Claude Opus 5.5May 2026not published2026-09-22
99%
Claude Sonnet 5.5May 2026not published2026-09-28
99%
GPT-6 SolApr 2026not published2026-09-22
100%
GPT-6 Luna ProApr 2026not published2026-09-22
98%
GPT-6 LunaApr 2026not published2026-09-22
99%
GPT-6.1 Sol ProMar 2026not published2026-09-29
98%
GPT-6.1 SolMar 2026not published2026-09-29
98%

Each bar is one month from Jan 2024; height = share answered correctly; pink = measured horizon.

How it works

  1. For each month, a writing model drafts questions from Wikipedia's Current events portal. A question is kept only if its answer appears word for word in the source event line and not in the question.
  2. Eight questions per month from January 2024. A new month is added automatically after it ends.
  3. Models answer with a short phrase or "unknown", temperature 0, lowest reasoning setting, no tools. Grading is string matching against the answer and its aliases.
  4. Horizon = the last month whose three-month rolling accuracy is at least half of the model's 2024 accuracy.
Open dataset · v1

264 items, every model answer and every score are public under CC BY 4.0. Cite as “BenchGecko Gecko Tests, Knowledge Horizon v1” with a link.

Download JSON
Training data thins out in the final months before a cutoff, and some labs publish only a year. The measured horizon shows what the model can actually use.