1 paper · 1 filter
Kentaro Mitsui, Koh Mitsuda, Toshiaki Wakatsuki +2
Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in resp…