BIG-Bench Hard

23 BIG-Bench tasks where models initially performed below the average human rater, focused on multi-step reasoning.