3 citations · 4 across the 13 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2024
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment
Chenliang Li, Siliang Zeng, Zeyi Liao +4
Aligning human preference and value is an important requirement for building contemporary foundation models and embodied AI. However, popular approaches such as reinforcement learn…
cs.AI2024★ 1 cited
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
Jiaxiang Li, Siliang Zeng, Hoi-To Wai +3
Aligning human preference and value is an important requirement for contemporary foundation models. State-of-the-art techniques such as Reinforcement Learning from Human Feedback (…