Terminal-Bench-Science: Evaluating AI agents on scientific research workflowsBy matt_d · Discussion · HNStory 49472820 · Front PageView original