9 papers · 1 filter
DR-Arena: an Automated Evaluation Framework for Deep Research Agents
Yiwen Gao, Ruochen Zhao, Yang Deng +1
As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, reliable evaluation of their task p…
BranPO: Scalable Contrastive Branch Sampling for Long-Horizon Agentic Reinforcement Learning
Yubao Zhao, Weiquan Huang, Sudong Wang +4
Agentic reinforcement learning enables large language models to perform multi-turn planning and tool use, but long-horizon training remains challenging under sparse trajectory-leve…
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
Songjun Tu, Yiwen Ma, Jiahao Lin +6
Large language models can generate fluent peer reviews, yet their assessments often lack sufficient critical rigor when substantive issues are subtle and distributed across a paper…
A Comprehensive Survey of Contamination Detection Methods in Large Language Models
Mathieu Ravaut, Bosheng Ding, Fangkai Jiao +6
With the rise of Large Language Models (LLMs) in recent years, abundant new opportunities are emerging, but also new challenges, among which contamination is quickly becoming criti…
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia +3
As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly. Currently, static benchmarks suffer from inflexibi…
Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
Xingxuan Li, Weiwen Xu, Ruochen Zhao +3
State-of-the-art large language models (LLMs) exhibit impressive problem-solving capabilities but may struggle with complex reasoning and factual correctness. Existing methods harn…