collaborators

9 papers

cs.AI2026

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Zhilin Wang, Han Song, Runzhe Zhan +13

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it…

cs.AI2026

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

Xinyang Liao, Lingyu Li, Huacan Liu +5

As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly important. Otherwise, such systems may rapidly…

cs.AI2026

-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

Haoran Zhang, Luxin Xu, Zhilin Wang +11

The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life and work. A core challenge in…

cs.AI2026

Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling

Yafu Li, Runzhe Zhan, Haoran Zhang +25

Recent progress in reasoning models has substantially advanced long-horizon mathematical and scientific problem solving, with several systems now reaching gold-medal-level performa…

cs.GT2026

Common-agency Games for Multi-Objective Test-Time Alignment

Baiting Chen, Tong Zhu, Rui Yu +1

Aligning large language models (LLMs) with human preferences is inherently multi-objective: different users and evaluation criteria impose heterogeneous and often conflicting requi…

cs.LG2026

ALIGN: Aligned Delegation with Performance Guarantees for Multi-Agent LLM Reasoning

Tong Zhu, Baiting Chen, Jin Zhou +3

LLMs often underperform on complex reasoning tasks when relying on a single generation-and-selection pipeline. Inference-time ensemble methods can improve performance by sampling d…