collaborators

19 papers

cs.AI2026

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Zhilin Wang, Han Song, Runzhe Zhan +13

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it…

cs.LG2026

How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs

Zhichen Dong, Yang Li, Yuhan Sun +9

Token-level credit assignment remains a key obstacle for reinforcement learning (RL) in large language models (LLMs), where RL recipes typically treat all tokens equally, failing t…

cs.AI2026

ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics

Shunkai Zhang, Haoran Zhang, Yun Luo +15

Combinatorics is central to Olympiad-level mathematical problem solving, requiring deep discrete reasoning, creative constructions, and rigorous structural insight. Recent evidence…

cs.AI2026

SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents

Wenxuan Wang, Haoyu Sun, Fukuan Hou +4

Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions. As these memories grow, they may reinforce one another, di…

cs.CL2026

Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

Haoyu Sun, Wenxuan Wang, Mingyang Song +5

Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible. Yet existing agent eva…

cs.CL2026

Characterizing, Evaluating, and Optimizing Complex Reasoning

Haoran Zhang, Yafu Li, Zhi Wang +4

Large Reasoning Models (LRMs) increasingly rely on reasoning traces with complex internal structures. However, existing work lacks a unified answer to three fundamental questions:…