7 papers
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling
Evan Assmus, Qining Zhang, Lei Ying
Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. Existing reward-model-free methods are either restricted to bandi…
Multi-Metric Adaptive Experimental Design Under a Fixed Budget with Validation
Qining Zhang, Tanner Fiez, Yi Liu +1
A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inf…
Efficient Federated RLHF via Zeroth-Order Policy Optimization
Deyi Wang, Qining Zhang, Lei Ying
This paper considers reinforcement learning from human feedback in a federated learning setting with resource-constrained agents, such as edge devices. We propose an efficient fede…
Early Lung Cancer Diagnosis from Virtual Follow-up LDCT Generation via Correlational Autoencoder and Latent Flow Matching
Yutong Wu, Yifan Wang, Qining Zhang +2
Lung cancer is one of the most commonly diagnosed cancers, and early diagnosis is critical because the survival rate declines sharply once the disease progresses to advanced stages…
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function
Qining Zhang, Lei Ying
The link function, which characterizes the relationship between the preference for two trajectories and their returns, is a crucial component in designing RL algorithms that learn…
Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Qining Zhang, Lei Ying
Reward inference (learning a reward model from human preferences) is a critical intermediate step in the Reinforcement Learning from Human Feedback (RLHF) pipeline for fine-tuning…