18 citations · 25 across the 19 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
Xiaobo Wang, Zixia Jia, Jiaqi Li +2
Offline preference optimization methods are efficient for large language models (LLMs) alignment. Direct Preference optimization (DPO)-like learning, one of the most popular approa…
cs.LG2024★ 1 cited
Tree-based Ensemble Learning for Out-of-distribution Detection
Zhaiming Shen, Menglun Wang, Guang Cheng +5
Being able to successfully determine whether the testing samples has similar distribution as the training samples is a fundamental question to address before we can safely deploy m…