10 papers
Resilience Beyond Pairwise Networks
Amitosh Tiwari, Chittaranjan Hens, Prosenjit Kundu
We derive a one-dimensional reduction for nonlinear dynamics on simplicial complexes containing both pairwise and triangular (higher-order) interactions. The effective state is def…
To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning
Fengji Zhang, Tianyu Fan, Yuxiang Zheng +4
Recent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved new state-of-the-art results on open-domain…
CLI-Anything: Towards Agent-Native Computer Use
Yuhao Yang, Tianyu Fan, Chao Huang
As large language models advance in reasoning and tool use capabilities, researchers increasingly seek to leverage them for computer use agents that can interact with existing soft…
Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory
Yuhao Zhan, Tianyu Fan, Linxuan Huang +2
Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evaluation, obscuring intermediate hal…
DeepInnovator: Triggering the Innovative Capabilities of LLMs
Tianyu Fan, Fengji Zhang, Yuxiang Zheng +5
The application of Large Language Models (LLMs) in accelerating scientific discovery has garnered increasing attention, with a key focus on constructing research agents endowed wit…
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
Yumeng Wang, Tianyu Fan, Lingrui Xu +1
Large Language Models (LLMs) have evolved from simple chatbots into sophisticated agents capable of automating complex real-world tasks, where browsing and reasoning over live web…