17 papers
Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees
Yu Chen, Ruishuo Chen, Xun Wang +2
Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill select…
MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models
Chenglin Liu, Xun Wang, Ruishuo Chen +2
Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward generation, existing methods generate and re…
When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict
Lu Yang, Shusheng Xu, Zhuoran Li +2
LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks…
When Context Returns: Toward Robust Internalization in On-Policy Distillation
Xun Wang, Ruishuo Chen, Zhuoran Li +2
Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so that the context is no longer ne…
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
Ruishuo Chen, Yu Chen, Zhuoran Li +1
Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without…
Reparameterization Flow Policy Optimization
Hai Zhong, Zhuoran Li, Xun Wang +1
Reparameterization Policy Gradient (RPG) has emerged as a powerful paradigm for model-based reinforcement learning, enabling high sample efficiency by backpropagating gradients thr…