activity
20182026
most citedFederated Multi-armed Bandits with Personalization

15 citations · 52 across the 21 of their papers we have counts for

collaborators
Showing cs.LGShow all

14 papers · 1 filter

cs.LG2026

-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses

Di Wu, Chengshuai Shi, Jing Yang +1

Reinforcement Learning from Human Feedback (RLHF) has become a cornerstone technique for post-training large language models. While most existing approaches rely on the reverse KL-…

cs.LG2026

Lean Clients, Full Accuracy: Hybrid Zeroth- and First-Order Split Federated Learning

Zhoubin Kou, Zihan Chen, Jing Yang +1

Split Federated Learning (SFL) enables collaborative training between resource-constrained edge devices and a compute-rich server. Communication overhead is a central issue in SFL…

cs.LG2025

Greedy Sampling Is Provably Efficient for RLHF

Di Wu, Chengshuai Shi, Jing Yang +1

Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique for post-training large language models. Despite its empirical success, the theoretical understandi…

cs.LG2025

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

Ruiquan Huang, Donghao Li, Chengshuai Shi +2

This paper investigates a hybrid learning framework for reinforcement learning (RL) in which the agent can leverage both an offline dataset and online interactions to learn the opt…

cs.LG2025

Chain-of-Thought Enhanced Shallow Transformers for Wireless Symbol Detection

Li Fan, Peng Wang, Jing Yang +1

Transformers have shown potential in solving wireless communication problems, particularly via in-context learning (ICL), where models adapt to new tasks through prompts without re…

cs.LG2025

A Shared Low-Rank Adaptation Approach to Personalized RLHF

Renpu Liu, Peng Wang, Donghao Li +2

Reinforcement Learning from Human Feedback (RLHF) has emerged as a pivotal technique for aligning artificial intelligence systems with human values, achieving remarkable success in…