4 papers
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
Qiyao Ma, Dechen Gao, Rui Cai +4
Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing d…
VITA: Vision-to-Action Flow Matching Policy
Dechen Gao, Boqi Zhao, Andrew Lee +6
Conventional flow matching and diffusion-based policies sample via iterative denoising from standard noise distributions (e.g., Gaussian), and require conditioning modules to repea…
IN-RIL: Interleaved Reinforcement and Imitation Learning for Policy Fine-Tuning
Dechen Gao, Hang Wang, Hanchu Zhou +5
Imitation learning (IL) and reinforcement learning (RL) each offer distinct advantages for robotics policy learning: IL provides stable learning from demonstrations, and RL promote…
EI-Drive: A Platform for Cooperative Perception with Realistic Communication Models
Hanchu Zhou, Edward Xie, Wei Shao +3
The growing interest in autonomous driving calls for realistic simulation platforms capable of accurately simulating cooperative perception process in realistic traffic scenarios.…