Evals

Benchmarks and leaderboards.

MMLU Benchmark

Massive Multitask Language Understanding — 57-subject benchmark covering STEM, humanities, and more

benchmark

HumanEval

164 handwritten Python problems for evaluating LLM code generation ability

benchmark