2 papers
cs.LG2026
Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
Zhu Zhang, Jixun Wang, Xiaoang Xu +6
On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher sup…
cs.CL2026
Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
Yuzhi Tang, Wentao Ma, Xiling Zhao +17
Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-taking behavior when explicitly instructed…