1 paper
Ziqi Wang, Hongshuo Huang, Hancheng Zhao +4
Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents lose context, game verification checks,…