Multi-turn conversation quality judged by a strong model acting as grader across eight categories.
Read the original source — github.com
benchmark · Shared by tscosj
0 comments
No comments yet.
Multi-turn conversation quality judged by a strong model acting as grader across eight categories.
Multi-turn conversation quality judged by a strong model acting as grader across eight categories.
Read the original source — github.com
benchmark · Shared by tscosj
No comments yet.