activity
20242026
collaborators

11 papers

cs.RO2026

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

Baiqi Li, Ce Zhang, Yu Fang +4

A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that…

cs.LG2026

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning

Shangzhe Li, Xuchao Zhang, Chetan Bansal +1

Self-play post-training methods has emerged as an effective approach for finetuning large language models and turn the weak language model into strong language model without prefer…

cs.LG2026

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation

Shangzhe Li, Weitong Zhang

We study value adaptation in offline-to-online reinforcement learning under general function approximation. Starting from an imperfect offline pretrained -function, the learner…

cs.LG2026

Quantile Q-Learning: Revisiting Offline Extreme Q-Learning with Quantile Regression

Xinming Gao, Shangzhe Li, Yujin Cai +1

Offline reinforcement learning (RL) enables policy learning from fixed datasets without further environment interaction, making it particularly valuable in high-risk or costly doma…

cs.LG2026

Imitation from Observations with Trajectory-Level Generative Embeddings

Yongtao Qu, Shangzhe Li, Weitong Zhang

We consider the offline imitation learning from observations (LfO) where the expert demonstrations are scarce and the available offline suboptimal data are far from the expert beha…

cs.LG2026

Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning

Shangzhe Li, Dongruo Zhou, Weitong Zhang

We study online adversarial imitation learning (AIL), where an agent learns from offline expert demonstrations and interacts with the environment online without access to rewards.…