Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning
Qingjun Wang, Hongtu Zhou, Hang Yu +5
Offline reinforcement learning (RL) faces a critical challenge of overestimating the value of out-of-distribution (OOD) actions. Existing methods mitigate this issue by penalizing…
cs.LG2025
On Weak-to-Strong Generalization and f-Divergence
Wei Yao, Gengze Xu, Huayi Tang +4
Weak-to-strong generalization (W2SG) has emerged as a promising paradigm for stimulating the capabilities of strong pre-trained models by leveraging supervision from weaker supervi…
cs.LG2025
Weak-to-Strong Generalization via Bregman Bias-Variance Decomposition
Gengze Xu, Wei Yao, Ziqiao Wang +1
Weak-to-strong generalization (W2SG) is the phenomenon in which a powerful student model, trained on labels produced by a weaker teacher, ultimately outperforms the teacher on the…