Head-to-head instruction-following judged by an LLM, reported as a win rate against a reference model.
Read the original source — tatsu-lab.github.io
benchmark · Shared by tscosj
0 comments
No comments yet.
Head-to-head instruction-following judged by an LLM, reported as a win rate against a reference model.
Head-to-head instruction-following judged by an LLM, reported as a win rate against a reference model.
Read the original source — tatsu-lab.github.io
benchmark · Shared by tscosj
No comments yet.