End-to-end tasks completed in a real terminal sandbox, testing long-horizon agentic execution.
Read the original source — tbench.ai
benchmark · Shared by tscosj
0 comments
No comments yet.
End-to-end tasks completed in a real terminal sandbox, testing long-horizon agentic execution.
End-to-end tasks completed in a real terminal sandbox, testing long-horizon agentic execution.
Read the original source — tbench.ai
benchmark · Shared by tscosj
No comments yet.