15 citations · 52 across the 21 of their papers we have counts for
14 papers · 1 filter
-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
Di Wu, Chengshuai Shi, Jing Yang +1
Reinforcement Learning from Human Feedback (RLHF) has become a cornerstone technique for post-training large language models. While most existing approaches rely on the reverse KL-…
Lean Clients, Full Accuracy: Hybrid Zeroth- and First-Order Split Federated Learning
Zhoubin Kou, Zihan Chen, Jing Yang +1
Split Federated Learning (SFL) enables collaborative training between resource-constrained edge devices and a compute-rich server. Communication overhead is a central issue in SFL…
Greedy Sampling Is Provably Efficient for RLHF
Di Wu, Chengshuai Shi, Jing Yang +1
Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique for post-training large language models. Despite its empirical success, the theoretical understandi…
Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis
Ruiquan Huang, Donghao Li, Chengshuai Shi +2
This paper investigates a hybrid learning framework for reinforcement learning (RL) in which the agent can leverage both an offline dataset and online interactions to learn the opt…
Chain-of-Thought Enhanced Shallow Transformers for Wireless Symbol Detection
Li Fan, Peng Wang, Jing Yang +1
Transformers have shown potential in solving wireless communication problems, particularly via in-context learning (ICL), where models adapt to new tasks through prompts without re…
A Shared Low-Rank Adaptation Approach to Personalized RLHF
Renpu Liu, Peng Wang, Donghao Li +2
Reinforcement Learning from Human Feedback (RLHF) has emerged as a pivotal technique for aligning artificial intelligence systems with human values, achieving remarkable success in…