10 papers
Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees
Yu Chen, Ruishuo Chen, Xun Wang +2
Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill select…
MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models
Chenglin Liu, Xun Wang, Ruishuo Chen +2
Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward generation, existing methods generate and re…
When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict
Lu Yang, Shusheng Xu, Zhuoran Li +2
LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks…
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Lingao Xiao, Yalun Dai, Yangyu Huang +17
Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. Recent systems increasingly generate multi…
Reparameterization Flow Policy Optimization
Hai Zhong, Zhuoran Li, Xun Wang +1
Reparameterization Policy Gradient (RPG) has emerged as a powerful paradigm for model-based reinforcement learning, enabling high sample efficiency by backpropagating gradients thr…
From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos
Xun Wang, Zhuoran Li, Yanshan Lin +2
Training a team of agents from scratch in multi-agent reinforcement learning (MARL) is highly inefficient, much like asking beginners to play a symphony together without first prac…