1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Shuaiyi Huang, Mara Levy, Anubhav Gupta +3
Preference feedback collected by human or VLM annotators is often noisy, presenting a significant challenge for preference-based reinforcement learning that relies on accurate pref…