8 papers
Weak-to-Strong Generalization via Direct On-Policy Distillation
Shiyuan Feng, Huan-ang Gao, Haohan Chi +7
Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because t…
Provably Efficient Policy-Reward Co-Pretraining for Adversarial Imitation Learning
Tian Xu, Zexuan Chen, Zhilong Zhang +4
Adversarial imitation learning (AIL) achieves high-quality imitation compared to behavioral cloning (BC), but demands substantial online environment interaction. Recent empirical w…
Planning with Unified Multimodal Models
Yihao Sun, Zhilong Zhang, Yang Yu +1
With the powerful reasoning capabilities of large language models (LLMs) and vision-language models (VLMs), many recent works have explored using them for decision-making. However,…
Continual Quadruped Robots Coordination via Semantic Skill Discovery
Daoqing Wang, Yuchen Xiao, Weixuan Huang +5
Multi-quadruped coordination has attracted increasing attention due to its enhanced payload capacity, broader contact coverage, and improved adaptability to challenging tasks. Exis…
Adversarial Imitation Learning with General Function Approximation: Theoretical Analysis and Practical Algorithms
Tian Xu, Zhilong Zhang, Zexuan Chen +3
Adversarial imitation learning (AIL), a prominent approach in imitation learning, has achieved significant practical success powered by neural network approximation. However, exist…
Anticipation-VLA: Solving Long-Horizon Embodied Tasks via Anticipation-based Subgoal Generation
Zhilong Zhang, Wenyu Luo, Haonan Wang +9
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, enabling robots to perform tasks based on natural language instructions and curre…