collaborators

7 papers

cs.LG2026

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

Evan Assmus, Qining Zhang, Lei Ying

Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. Existing reward-model-free methods are either restricted to bandi…

cs.LG2026

Multi-Metric Adaptive Experimental Design Under a Fixed Budget with Validation

Qining Zhang, Tanner Fiez, Yi Liu +1

A/B tests in online experiments face statistical power challenges when testing multiple candidates simultaneously, while adaptive experimental designs (AED) alone fall short in inf…

cs.LG2026

Efficient Federated RLHF via Zeroth-Order Policy Optimization

Deyi Wang, Qining Zhang, Lei Ying

This paper considers reinforcement learning from human feedback in a federated learning setting with resource-constrained agents, such as edge devices. We propose an efficient fede…

cs.CV2025

Early Lung Cancer Diagnosis from Virtual Follow-up LDCT Generation via Correlational Autoencoder and Latent Flow Matching

Yutong Wu, Yifan Wang, Qining Zhang +2

Lung cancer is one of the most commonly diagnosed cancers, and early diagnosis is critical because the survival rate declines sharply once the disease progresses to advanced stages…

cs.LG2025

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

Qining Zhang, Lei Ying

The link function, which characterizes the relationship between the preference for two trajectories and their returns, is a crucial component in designing RL algorithms that learn…

cs.LG2025

Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference

Qining Zhang, Lei Ying

Reward inference (learning a reward model from human preferences) is a critical intermediate step in the Reinforcement Learning from Human Feedback (RLHF) pipeline for fine-tuning…