1 citations · 1 across the 7 of their papers we have counts for
4 papers · 1 filter
On-policy Distillation with Verifiable Reward
Wenze Lin, Jiale Zhao, Xitai Jiang +5
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RL…
Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual
Yining Li, Peizhong Ju, Ness Shroff
Reinforcement Learning from Human Feedback (RLHF) plays a significant role in aligning Large Language Models (LLMs) with human preferences. While RLHF with expected reward constrai…
How to Find the Exact Pareto Front for Multi-Objective MDPs?
Yining Li, Peizhong Ju, Ness B. Shroff
Multi-Objective Markov Decision Processes (MO-MDPs) are receiving increasing attention, as real-world decision-making problems often involve conflicting objectives that cannot be a…
Achieving Sample and Computational Efficient Reinforcement Learning by Action Space Reduction via Grouping
Yining Li, Peizhong Ju, Ness Shroff
Reinforcement learning often needs to deal with the exponential growth of states and actions when exploring optimal control in high-dimensional spaces (often known as the curse of…