2 papers
cs.AI2026
Building Agent Harnesses for Scientific Curation from Multimodal Sources
Sheng Zhang, Qin Liu, Renqian Luo +9
Scientific discovery workflows often depend on structured curation from the literature. This is difficult for current agents because the key evidence is scattered across long text,…
cs.AI2026
OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents
Kaicheng Zhang, Wen Ge, Lei Jiang +5
Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tasks, while the financial rele…