2 papers
cs.CL2026
Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science
Maohao Ran, Chendong Ma, Yanting Zhang +4
Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unob…
cs.AI2026
CaveAgent: Transforming LLMs into Stateful Runtime Operators
Maohao Ran, Zhenglin Wan, Cooper Lin +21
LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradigms that struggle with long-horizon tasks…