collaborators

5 papers

cs.LG2026

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

Evan Assmus, Qining Zhang, Lei Ying

Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. Existing reward-model-free methods are either restricted to bandi…

cs.LG2026

Efficient Federated RLHF via Zeroth-Order Policy Optimization

Deyi Wang, Qining Zhang, Lei Ying

This paper considers reinforcement learning from human feedback in a federated learning setting with resource-constrained agents, such as edge devices. We propose an efficient fede…

cs.LG2025

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

Qining Zhang, Lei Ying

The link function, which characterizes the relationship between the preference for two trajectories and their returns, is a crucial component in designing RL algorithms that learn…

cs.LG2025

Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference

Qining Zhang, Lei Ying

Reward inference (learning a reward model from human preferences) is a critical intermediate step in the Reinforcement Learning from Human Feedback (RLHF) pipeline for fine-tuning…

cs.LG2025

Reinforcement Learning from Human Feedback without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis

Qining Zhang, Honghao Wei, Lei Ying

In this paper, we study reinforcement learning from human feedback (RLHF) under an episodic Markov decision process with a general trajectory-wise reward model. We developed a mode…