2 papers
cs.SD2026
Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
Fangxu Yu, Tao Feng, Dehai Min +6
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs…
cs.LG2026
Weak-to-Strong On-Policy Distillation
Fangxu Yu, Weijia Xu, Michael Xu +2
On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabiliti…