6 papers
DR-Arena: an Automated Evaluation Framework for Deep Research Agents
Yiwen Gao, Ruochen Zhao, Yang Deng +1
As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, reliable evaluation of their task p…
BranPO: Scalable Contrastive Branch Sampling for Long-Horizon Agentic Reinforcement Learning
Yubao Zhao, Weiquan Huang, Sudong Wang +4
Agentic reinforcement learning enables large language models to perform multi-turn planning and tool use, but long-horizon training remains challenging under sparse trajectory-leve…
Unified-MAS: Universally Generating Domain-Specific Nodes for Empowering Automatic Multi-Agent Systems
Hehai Lin, Yu Yan, Zixuan Wang +6
Automatic Multi-Agent Systems (MAS) generation has emerged as a promising paradigm for solving complex reasoning tasks. However, existing frameworks are fundamentally bottlenecked…
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
Songjun Tu, Yiwen Ma, Jiahao Lin +6
Large language models can generate fluent peer reviews, yet their assessments often lack sufficient critical rigor when substantive issues are subtle and distributed across a paper…
AgREE: Agentic Reasoning for Knowledge Graph Completion on Emerging Entities
Ruochen Zhao, Simone Conia, Eric Peng +2
Open-domain Knowledge Graph Completion (KGC) faces significant challenges in an ever-changing world, especially when considering the continual emergence of new entities in daily ne…
A Comprehensive Survey of Contamination Detection Methods in Large Language Models
Mathieu Ravaut, Bosheng Ding, Fangkai Jiao +6
With the rise of Large Language Models (LLMs) in recent years, abundant new opportunities are emerging, but also new challenges, among which contamination is quickly becoming criti…