collaborators

8 papers

cs.LG2026

Weak-to-Strong Generalization via Direct On-Policy Distillation

Shiyuan Feng, Huan-ang Gao, Haohan Chi +7

Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because t…

cs.LG2026

Provably Efficient Policy-Reward Co-Pretraining for Adversarial Imitation Learning

Tian Xu, Zexuan Chen, Zhilong Zhang +4

Adversarial imitation learning (AIL) achieves high-quality imitation compared to behavioral cloning (BC), but demands substantial online environment interaction. Recent empirical w…

cs.CV2026

Planning with Unified Multimodal Models

Yihao Sun, Zhilong Zhang, Yang Yu +1

With the powerful reasoning capabilities of large language models (LLMs) and vision-language models (VLMs), many recent works have explored using them for decision-making. However,…

cs.RO2026

Continual Quadruped Robots Coordination via Semantic Skill Discovery

Daoqing Wang, Yuchen Xiao, Weixuan Huang +5

Multi-quadruped coordination has attracted increasing attention due to its enhanced payload capacity, broader contact coverage, and improved adaptability to challenging tasks. Exis…

cs.LG2026

Adversarial Imitation Learning with General Function Approximation: Theoretical Analysis and Practical Algorithms

Tian Xu, Zhilong Zhang, Zexuan Chen +3

Adversarial imitation learning (AIL), a prominent approach in imitation learning, has achieved significant practical success powered by neural network approximation. However, exist…

cs.RO2026

Anticipation-VLA: Solving Long-Horizon Embodied Tasks via Anticipation-based Subgoal Generation

Zhilong Zhang, Wenyu Luo, Haonan Wang +9

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, enabling robots to perform tasks based on natural language instructions and curre…