13 papers
A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination
Wenxiao Zhao, Dong Liu, Kaiyi Xu +10
Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search fai…
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
Haobo Li, Eunseo Jung, Wenxiao Zhao +8
The paper presents AutoSupervision, a system that automatically verifies whether manuscript revisions truly address reviewer comments by grounding the verification in evidence from…
Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
Chuhan Shi, Xiaoquan Ren, Sicheng Song +3
The paper presents SDABench, a capability-oriented benchmark that evaluates large language models on scientific data analysis tasks across biology, chemistry, environment, geograph…
WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning
Zixin Chen, Peng Liu, Haobo Li +7
Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report,…
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao +9
Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories, requiring them to locate r…
TeachArena: Are Language Agents Ready for Realistic Teaching Work?
Zixin Chen, Peng Liu, Rui Sheng +6
Language agents are increasingly deployed in professional workflows, yet tutoring remains a high-stakes capability that existing evaluations only partially capture. Effective tutor…