activity
20202026
most citedDynamic Knapsack Optimization Towards Efficient Multi-Channel Sequential Advertising

9 citations · 16 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Jing Liang, Hongyao Tang, Yi Ma +9

Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse. O…

cs.LG2026

The Rank and Gradient Lost in Non-stationarity: Sample Weight Decay for Mitigating Plasticity Loss in Reinforcement Learning

Zihao Wu, Hongyao Tang, Yi Ma +3

Deep reinforcement learning (RL) suffers from plasticity loss severely due to the nature of non-stationarity, which impairs the ability to adapt to new data and learn continually.…

cs.LG20243 cited

Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback

Yifu Yuan, Jianye Hao, Yi Ma +6

Reinforcement Learning with Human Feedback (RLHF) has received significant attention for performing tasks without the need for costly manual reward design by aligning human prefere…

cs.LG20234 cited

Rethinking Decision Transformer via Hierarchical Reinforcement Learning

Yi Ma, Chenjun Xiao, Hebin Liang +1

Decision Transformer (DT) is an innovative algorithm leveraging recent advances of the transformer architecture in reinforcement learning (RL). However, a notable limitation of DT…

cs.LG2023

Scaff-PD: Communication Efficient Fair and Robust Federated Learning

Yaodong Yu, Sai Praneeth Karimireddy, Yi Ma +1

We present Scaff-PD, a fast and communication-efficient algorithm for distributionally robust federated learning. Our approach improves fairness by optimizing a family of distribut…

cs.LG2022

State-Aware Proximal Pessimistic Algorithms for Offline Reinforcement Learning

Chen Chen, Hongyao Tang, Yi Ma +4

Pessimism is of great importance in offline reinforcement learning (RL). One broad category of offline RL algorithms fulfills pessimism by explicit or implicit behavior regularizat…