WildBench

Real-world user prompts drawn from chats, graded against reference answers by an LLM judge.