collaborators

10 papers

cs.LG2026

-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses

Di Wu, Chengshuai Shi, Jing Yang +1

Reinforcement Learning from Human Feedback (RLHF) has become a cornerstone technique for post-training large language models. While most existing approaches rely on the reverse KL-…

cs.LG2026

Lean Clients, Full Accuracy: Hybrid Zeroth- and First-Order Split Federated Learning

Zhoubin Kou, Zihan Chen, Jing Yang +1

Split Federated Learning (SFL) enables collaborative training between resource-constrained edge devices and a compute-rich server. Communication overhead is a central issue in SFL…

cs.LG2025

On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery

Renpu Liu, Ruida Zhou, Cong Shen +1

An intriguing property of the Transformer is its ability to perform in-context learning (ICL), where the Transformer can solve different inference tasks without parameter updating…

cs.LG2025

Greedy Sampling Is Provably Efficient for RLHF

Di Wu, Chengshuai Shi, Jing Yang +1

Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique for post-training large language models. Despite its empirical success, the theoretical understandi…

cs.IT2025

Decision Feedback In-Context Learning for Wireless Symbol Detection

Li Fan, Wei Shen, Jing Yang +1

Pre-trained Transformers, through in-context learning (ICL), have demonstrated exceptional capabilities to adapt to new tasks using example prompts without model update. Transforme…

cs.LG2025

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

Ruiquan Huang, Donghao Li, Chengshuai Shi +2

This paper investigates a hybrid learning framework for reinforcement learning (RL) in which the agent can leverage both an offline dataset and online interactions to learn the opt…