2 papers
cs.AI2026
SteerDuplex: Steerable Duplex Speech Dialogue Models
Utkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai +13
Full-duplex spoken dialogue models support low-latency turn taking, interruption handling, and backchanneling, yet a key capability remains underexplored: steerability, the ability…
cs.CL2025
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Jeff Da, Clinton Wang, Xiang Deng +3
Reinforcement Learning from Verifiable Rewards (RLVR) has been widely adopted as the de facto method for enhancing the reasoning capabilities of large language models and has demon…