activity
20242026
collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

DR-Arena: an Automated Evaluation Framework for Deep Research Agents

Yiwen Gao, Ruochen Zhao, Yang Deng +1

As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, reliable evaluation of their task p…

cs.CL2026

BranPO: Scalable Contrastive Branch Sampling for Long-Horizon Agentic Reinforcement Learning

Yubao Zhao, Weiquan Huang, Sudong Wang +4

Agentic reinforcement learning enables large language models to perform multi-turn planning and tool use, but long-horizon training remains challenging under sparse trajectory-leve…

cs.CL20262 cited

PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review

Songjun Tu, Yiwen Ma, Jiahao Lin +6

Large language models can generate fluent peer reviews, yet their assessments often lack sufficient critical rigor when substantive issues are subtle and distributed across a paper…

cs.CL2025

A Comprehensive Survey of Contamination Detection Methods in Large Language Models

Mathieu Ravaut, Bosheng Ding, Fangkai Jiao +6

With the rise of Large Language Models (LLMs) in recent years, abundant new opportunities are emerging, but also new challenges, among which contamination is quickly becoming criti…

cs.CL2024

Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions

Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia +3

As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly. Currently, static benchmarks suffer from inflexibi…

cs.CL2024

Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks

Xingxuan Li, Weiwen Xu, Ruochen Zhao +3

State-of-the-art large language models (LLMs) exhibit impressive problem-solving capabilities but may struggle with complex reasoning and factual correctness. Existing methods harn…