23 BIG-Bench tasks where models initially performed below the average human rater, focused on multi-step reasoning.
Read the original source — github.com
benchmark · Shared by tscosj
0 comments
No comments yet.
23 BIG-Bench tasks where models initially performed below the average human rater, focused on multi-step reasoning.
23 BIG-Bench tasks where models initially performed below the average human rater, focused on multi-step reasoning.
Read the original source — github.com
benchmark · Shared by tscosj
No comments yet.