Showing cs.LGShow all
3 papers · 1 filter
cs.LG2024
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
Junkang Wu, Yuexiang Xie, Zhengyi Yang +6
This study addresses the challenge of noise in training datasets for Direct Preference Optimization (DPO), a method for aligning Large Language Models (LLMs) with human preferences…
cs.LG2023
BSL: Understanding and Improving Softmax Loss for Recommendation
Junkang Wu, Jiawei Chen, Jiancan Wu +3
Loss functions steer the optimization direction of recommendation models and are critical to model performance, but have received relatively little attention in recent recommendati…
cs.LG2023★ 13 cited
Understanding Contrastive Learning via Distributionally Robust Optimization
Junkang Wu, Jiawei Chen, Jiancan Wu +3
This study reveals the inherent tolerance of contrastive learning (CL) towards sampling bias, wherein negative samples may encompass similar semantics (\eg labels). However, existi…