1 paper
Yadi Cao, Sicheng Lai, Jiahe Huang +12
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@…