Benchmarks, and their decay
Read · 7 minEvery model release leads with benchmark scores, and every benchmark has a shelf life. MMLU, once the headline test of broad knowledge, is now saturated — the best models score so highly that it no longer distinguishes them. The frontier has moved to harder, more specific measures: GPQA for graduate-level science, SWE-bench for resolving real software bugs, and composite indices that blend many tests. The skill is to ask what a benchmark actually measures, whether it can be gamed by training on similar data, and whether a one-point lead is signal or noise. Treat any single number the way you'd treat a single quarter's earnings: context first.