57 papers
Measuring the Gap Between Human and LLM Research Ideas
Ziyu Chen, Yilun Zhao, Arman Cohan
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how f…
MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks
Yanjun Shao, Xiangru Tang, Jiwoong Sohn +10
Complex medical reasoning requires integrating heterogeneous clinical evidence across multiple inference steps. Large language models (LLMs) now approach this through two routes: i…
VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding
Lin Fu, Zheyuan Yang, Yang Wang +3
We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding. It comprises 315K video reas…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
Alan Li, Yixin Liu, Arpan Sarkar +2
Scientific problem solving poses unique challenges for LLMs, requiring both deep domain knowledge and the ability to apply such knowledge through complex reasoning. While automated…
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
Yilun Zhao, Jinbiao Wei, Tingyu Song +3
Reasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity. This capability is increasingly important…