11 papers
WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation
Baiqi Li, Ce Zhang, Yu Fang +4
A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that…
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
Shangzhe Li, Xuchao Zhang, Chetan Bansal +1
Self-play post-training methods has emerged as an effective approach for finetuning large language models and turn the weak language model into strong language model without prefer…
Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
Shangzhe Li, Weitong Zhang
We study value adaptation in offline-to-online reinforcement learning under general function approximation. Starting from an imperfect offline pretrained -function, the learner…
Quantile Q-Learning: Revisiting Offline Extreme Q-Learning with Quantile Regression
Xinming Gao, Shangzhe Li, Yujin Cai +1
Offline reinforcement learning (RL) enables policy learning from fixed datasets without further environment interaction, making it particularly valuable in high-risk or costly doma…
Imitation from Observations with Trajectory-Level Generative Embeddings
Yongtao Qu, Shangzhe Li, Weitong Zhang
We consider the offline imitation learning from observations (LfO) where the expert demonstrations are scarce and the available offline suboptimal data are far from the expert beha…
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
Shangzhe Li, Dongruo Zhou, Weitong Zhang
We study online adversarial imitation learning (AIL), where an agent learns from offline expert demonstrations and interacts with the environment online without access to rewards.…