Hard prompts from Chatbot Arena, scored by an LLM judge against a fixed baseline for reproducibility.
Read the original source — github.com
benchmark · Shared by tscosj
0 comments
No comments yet.
Hard prompts from Chatbot Arena, scored by an LLM judge against a fixed baseline for reproducibility.
Hard prompts from Chatbot Arena, scored by an LLM judge against a fixed baseline for reproducibility.
Read the original source — github.com
benchmark · Shared by tscosj
No comments yet.