2 papers
cs.CL2026
Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
Manato Yaguchi, Yotaro Kubo, Hikaru Asano +1
Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying…
cs.LG2026
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
Ku Onoda, Paavo Parmas, Manato Yaguchi +1
In policy gradient reinforcement learning, access to a differentiable model enables 1st-order gradient estimation that accelerates learning compared to relying solely on derivative…