2 papers
cs.SD2026
Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
Fangxu Yu, Tao Feng, Dehai Min +6
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs…
cs.LG2026
Weak-to-Strong On-Policy Distillation
Fangxu Yu, Zinan Lin, Xiaodong Liu +4
The paper proposes Weak-to-Strong On-Policy Distillation (W2S-OPD), a method that improves a large language model by distilling knowledge from multiple weaker models using a constr…