papers
Publications (3)
cs.LG2026
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
Rasool Fakoor, Murdock Aubry, Nicholas Stranges +1
Reinforcement learning is structurally harder than supervised learning because the policy changes the data distribution it learns from. The resulting fragility is especially visibl…
cs.CL2026
What Is Missing: Interpretable Ratings for Large Language Model Outputs
Nicholas Stranges, Yimin Yang
Current Large Language Model (LLM) preference learning methods such as Proximal Policy Optimization and Direct Preference Optimization learn from direct rankings or numerical ratin…
cs.CL2026
Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
Yuzhi Tang, Wentao Ma, Xiling Zhao +17
Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-taking behavior when explicitly instructed…