activity
20232026
most citedTell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents

4 citations · 8 across the 27 of their papers we have counts for

collaborators
Showing 2026Show all

19 papers · 1 filter

cs.AI2026

StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

Yinghao Chen, Zixi Chen, Bingxiang He +7

Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. We argue that an ideal self-evolution method should share the…

cs.LG2026

On-policy Distillation with Verifiable Reward

Wenze Lin, Jiale Zhao, Xitai Jiang +5

Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RL…

cs.CL2026

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Deyao Hong, Yizhe Chi, Wenyi Li +7

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at b…

cs.AI2026

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Yizhe Chi, Wenyi Li, Deyao Hong +7

Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the t…

cs.LG2026

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Huan-ang Gao, Haohan Chi, Yong Yan +7

Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist s…

cs.AI2026

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Yuhao Zhan, Bingxiang He, Zecong Tang +1

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery afte…