1 citations · 1 across the 3 of their papers we have counts for
4 papers
Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation
Wenbo Zhang, Wenzhuo Zhou, Hengrui Cai +1
Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling…
MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving
Nan Yang, Zhanwen Liu, Linfeng Zhang +4
Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual token sequences, particularly i…
Bi-Level Offline Policy Optimization with Limited Exploration
Wenzhuo Zhou
We study offline reinforcement learning (RL) which seeks to learn a good policy based on a fixed, pre-collected dataset. A fundamental challenge behind this task is the distributio…
Stackelberg Batch Policy Learning
Wenzhuo Zhou, Annie Qu
Batch reinforcement learning (RL) defines the task of learning from a fixed batch of data lacking exhaustive exploration. Worst-case optimality algorithms, which calibrate a value-…