AlpacaEval

Head-to-head instruction-following judged by an LLM, reported as a win rate against a reference model.