4 papers
Understanding Diversity Collapse in RLVR via the Lens of Overtraining
Suqin Yuan, Jinkun Chen, Jiyang Zheng +6
Reinforcement learning with verifiable rewards (RLVR) has become a key approach for enhancing the reasoning abilities of large language models. However, RLVR often suffers from \em…
When More Experts Hurt: Underfitting in Multi-Expert Learning to Defer
Shuqi Liu, Yuzhou Cao, Lei Feng +2
Learning to Defer (L2D) enables a classifier to abstain from predictions and defer to an expert, and has recently been extended to multi-expert settings. In this work, we show that…
Establishing Linear Surrogate Regret Bounds for Convex Smooth Losses via Convolutional Fenchel-Young Losses
Yuzhou Cao, Han Bao, Lei Feng +1
Surrogate regret bounds, also known as excess risk bounds, bridge the gap between the convergence rates of surrogate and target losses. The regret transfer is lossless if the surro…
Understanding and Mitigating the Bias in Sample Selection for Learning with Noisy Labels
Qi Wei, Lei Feng, Haobo Wang +1
Learning with noisy labels aims to ensure model generalization given a label-corrupted training set. The sample selection strategy achieves promising performance by selecting a lab…