Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model
Qi Gou, Cam-Tu Nguyen
Large Language Models (LLMs) have become increasingly popular due to their ability to process and generate natural language. However, as they are trained on massive datasets of tex…
cs.CL2024
Momentum Posterior Regularization for Multi-hop Dense Retrieval
Zehua Xia, Yuyang Wu, Yiyun Xia +1
Multi-hop question answering (QA) often requires sequential retrieval (multi-hop retrieval), where each hop retrieves missing knowledge based on information from previous hops. To…
cs.CL2024
Reward Difference Optimization For Sample Reweighting In Offline RLHF
Shiqi Wang, Zhengze Zhang, Rui Zhao +2
With the rapid advances in Large Language Models (LLMs), aligning LLMs with human preferences become increasingly important. Although Reinforcement Learning with Human Feedback (RL…