3 papers
cs.LG2026
Unbiased Alignment for Large Language Models with Noisy Preferences
Jialiang Wang, Xianming Liu, Xiong Zhou +2
The alignment of large language models with human preferences is commonly achieved through Reinforcement Learning from Human Feedback or Direct Preference Optimization. However, th…
cs.LG2025
Variation-Bounded Loss for Noise-Tolerant Learning
Jialiang Wang, Xiong Zhou, Xianming Liu +4
Mitigating the negative impact of noisy labels has been aperennial issue in supervised learning. Robust loss functions have emerged as a prevalent solution to this problem. In this…
cs.LG2025
-Softmax: Approximating One-Hot Vectors for Mitigating Label Noise
Jialiang Wang, Xiong Zhou, Deming Zhai +3
Noisy labels pose a common challenge for training accurate deep neural networks. To mitigate label noise, prior studies have proposed various robust loss functions to achieve noise…