4 papers · 1 filter
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance
Kai Yan, Alexander G. Schwing, Yu-Xiong Wang
Reinforcement Learning with Verifiable Rewards (RLVR) has achieved great success in developing Large Language Models (LLMs) with chain-of-thought rollouts for many tasks such as ma…
Latent Wasserstein Adversarial Imitation Learning
Siqi Yang, Kai Yan, Alexander G. Schwing +1
Imitation Learning (IL) enables agents to mimic expert behavior by learning from demonstrations. However, traditional IL methods require large amounts of medium-to-high-quality dem…
Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision Transformers
Kai Yan, Alexander G. Schwing, Yu-Xiong Wang
Decision Transformers have recently emerged as a new and compelling paradigm for offline Reinforcement Learning (RL), completing a trajectory in an autoregressive way. While improv…
Offline Imitation from Observation via Primal Wasserstein State Occupancy Matching
Kai Yan, Alexander G. Schwing, Yu-xiong Wang
In real-world scenarios, arbitrary interactions with the environment can often be costly, and actions of expert demonstrations are not always available. To reduce the need for both…