MMLU Benchmark
Massive Multitask Language Understanding — 57-subject benchmark covering STEM, humanities, and more
benchmarkBenchmarks and leaderboards.
Massive Multitask Language Understanding — 57-subject benchmark covering STEM, humanities, and more
benchmarkOld benchmarks saturated — v2 replaces them with MMLU-Pro, GPQA, MuSR, BBH, IFEval, and MATH-Hard. All open models re-ranked. Several previously top models drop significantly.
newsStandardized safety evaluation across violent crimes, CSAM, weapons, self-harm, hate speech, election interference, and more. Tests system and user prompts in English and 7 other languages.
news