5 papers
CaveAgent: Transforming LLMs into Stateful Runtime Operators
Maohao Ran, Zhenglin Wan, Cooper Lin +21
LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradigms that struggle with long-horizon tasks…
Training Diffusion Policies via Prior-Mapping Co-Evolution
Chubin Zhang, Zhenglin Wan, Feng Chen +7
Reinforcement learning (RL) faces a persistent tension: policies that are stable to optimize (e.g., Gaussians) are often too simple to represent the multimodal action distributions…
Adversarial Dual On-Policy Distillation from Expressive Teacher
Zhenglin Wan, Jingxuan Wu, Xingrui Yu +5
Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradigm by modeling multi-modal e…
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
Zhenglin Wan, Jingxuan Wu, Xingrui Yu +4
Flow Matching (FM) has shown remarkable ability in modeling complex distributions and achieves strong performance in offline imitation learning for cloning expert behaviors. Howeve…
Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching
Jingxuan Wu, Zhenglin Wan, Xingrui Yu +4
Flow-based text-to-image models follow deterministic trajectories, making it costly to explore diverse modes under limited sampling budgets. Existing approaches to improving divers…