activity
20232026
collaborators

17 papers

cs.AI2026

Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees

Yu Chen, Ruishuo Chen, Xun Wang +2

Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill select…

cs.LG2026

MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

Chenglin Liu, Xun Wang, Ruishuo Chen +2

Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward generation, existing methods generate and re…

cs.AI2026

When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict

Lu Yang, Shusheng Xu, Zhuoran Li +2

LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks…

cs.LG2026

When Context Returns: Toward Robust Internalization in On-Policy Distillation

Xun Wang, Ruishuo Chen, Zhuoran Li +2

Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so that the context is no longer ne…

cs.CL2026

PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching

Ruishuo Chen, Yu Chen, Zhuoran Li +1

Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without…

cs.LG2026

Reparameterization Flow Policy Optimization

Hai Zhong, Zhuoran Li, Xun Wang +1

Reparameterization Policy Gradient (RPG) has emerged as a powerful paradigm for model-based reinforcement learning, enabling high sample efficiency by backpropagating gradients thr…