1 citations · 1 across the 1 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025★ 1 cited
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
Tong Yang, Bo Dai, Lin Xiao +1
Multi-agent reinforcement learning (MARL) lies at the heart of a plethora of applications involving the interaction of a group of agents in a shared unknown environment. A prominen…
cs.LG2024
Faster WIND: Accelerating Iterative Best-of- Distillation for LLM Alignment
Tong Yang, Jincheng Mei, Hanjun Dai +5
Recent advances in aligning large language models with human preferences have corroborated the growing importance of best-of-N distillation (BOND). However, the iterative BOND algo…
cs.LG2024
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Shicong Cen, Jincheng Mei, Katayoon Goshvadi +6
Reinforcement learning from human feedback (RLHF) has demonstrated great promise in aligning large language models (LLMs) with human preference. Depending on the availability of pr…