Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
TERMINAL-BENCH-SCIENCE
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
A benchmark for evaluating AI agents on research workflows across scientific domains
0 comments
No comments yet.