10 papers
Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning
Haoyuan Li, Zhengdong Hu, Jun Wang +2
This paper explores agentic 3D spatial understanding, i.e., MLLM agents performing 3D reasoning through tool use. Existing methods often misuse tools and exhibit biased tool prefer…
A Comprehensive Anatomy of Human and DeepSeek-R1 LLM Mathematical Reasoning
Yuxiang Chen, Jun Wang
The emergence of "Aha moments" in large language models, particularly DeepSeek-R1-0120, has raised the question of whether these systems genuinely reason or merely imitate the appe…
Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship
Yating Pan, Jiajun Zhang, Jun Wang +1
LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and quantitative signals. Humanities sc…
A Benchmark for Deep Information Synthesis
Debjit Paul, Daniel Murphy, Milan Gritta +14
Large language model (LLM)-based agents are increasingly used to solve complex tasks involving tool use, such as web browsing, code execution, and data analysis. However, current e…
Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents
SHengjie Ma, Chenlong Deng, Jiaxin Mao +5
While reinforcement learning (RL) enhances their ability to plan and reason across retrieval steps, we identify a critical failure mode in this setting: Tool-Call Hacking. Unlike e…
SAGE: Strategy-Adaptive Generation Engine for Query Rewriting
Teng Wang, Hailei Gong, Changwang Zhang +1
Query rewriting is pivotal for enhancing dense retrieval, yet current methods demand large-scale supervised data or suffer from inefficient reinforcement learning (RL) exploration.…