29 papers
Measuring the Gap Between Human and LLM Research Ideas
Ziyu Chen, Yilun Zhao, Arman Cohan
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how f…
MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks
Yanjun Shao, Xiangru Tang, Jiwoong Sohn +10
Complex medical reasoning requires integrating heterogeneous clinical evidence across multiple inference steps. Large language models (LLMs) now approach this through two routes: i…
Survey on Evaluation of LLM-based Agents
Asaf Yehudai, Lilach Eden, Alan Li +5
LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environments. This paper provides the f…
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
Tianyu Yang, Sihong Wu, Yilun Zhao +6
Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems involving both textual and visual modalities.…
ANCHOR: Branch-Point Data Generation for GUI Agents
Jinbiao Wei, Yilun Zhao, Kangqi Ni +1
End-to-end GUI agents for real desktop environments require large amounts of high-quality interaction data, yet collecting human demonstrations is expensive and existing synthetic…
AlphaResearch: Accelerating New Algorithm Discovery with Language Models
Zhaojian Yu, Kaiyue Feng, Yilun Zhao +3
LLMs have made significant progress in complex but easy-to-verify problems, yet they still struggle with discovering the unknown. In this paper, we present \textbf{AlphaResearch},…