2 papers
cs.AI2025
Rethinking the Design of Reinforcement Learning-Based Deep Research Agents
Yi Wan, Jiuqi Wang, Liam Li +3
Large language models (LLMs) augmented with external tools are increasingly deployed as deep research agents that gather, reason over, and synthesize web information to answer comp…
cs.LG2024
Reward Learning From Preference With Ties
Jinsong Liu, Dongdong Ge, Ruihao Zhu
Reward learning plays a pivotal role in Reinforcement Learning from Human Feedback (RLHF), ensuring the alignment of language models. The Bradley-Terry (BT) model stands as the pre…