9 papers
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
Zhilin Wang, Han Song, Runzhe Zhan +13
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it…
PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
Xinyang Liao, Lingyu Li, Huacan Liu +5
As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly important. Otherwise, such systems may rapidly…
-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
Haoran Zhang, Luxin Xu, Zhilin Wang +11
The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life and work. A core challenge in…
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
Yafu Li, Runzhe Zhan, Haoran Zhang +25
Recent progress in reasoning models has substantially advanced long-horizon mathematical and scientific problem solving, with several systems now reaching gold-medal-level performa…
Common-agency Games for Multi-Objective Test-Time Alignment
Baiting Chen, Tong Zhu, Rui Yu +1
Aligning large language models (LLMs) with human preferences is inherently multi-objective: different users and evaluation criteria impose heterogeneous and often conflicting requi…
ALIGN: Aligned Delegation with Performance Guarantees for Multi-Agent LLM Reasoning
Tong Zhu, Baiting Chen, Jin Zhou +3
LLMs often underperform on complex reasoning tasks when relying on a single generation-and-selection pipeline. Inference-time ensemble methods can improve performance by sampling d…