25 papers
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
Zhilin Wang, Han Song, Runzhe Zhan +13
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it…
ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics
Shunkai Zhang, Haoran Zhang, Yun Luo +15
Combinatorics is central to Olympiad-level mathematical problem solving, requiring deep discrete reasoning, creative constructions, and rigorous structural insight. Recent evidence…
From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG
Wenhao Wu, Zhentao Tang, Yafu Li +5
Large Language Models (LLMs) exhibit high reasoning capacity in medical question-answering, but their tendency to produce hallucinations and outdated knowledge poses critical risks…
Characterizing, Evaluating, and Optimizing Complex Reasoning
Haoran Zhang, Yafu Li, Zhi Wang +4
Large Reasoning Models (LRMs) increasingly rely on reasoning traces with complex internal structures. However, existing work lacks a unified answer to three fundamental questions:…
Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Deliberation
Haoran Zhang, Yafu Li, Xuyang Hu +4
Large language models (LLMs) are increasingly applied in diverse real-world scenarios, each governed by bespoke behavioral and safety specifications (spec) custom-tailored by users…
-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
Haoran Zhang, Luxin Xu, Zhilin Wang +11
The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life and work. A core challenge in…