1 paper · 1 filter
Jingjing Jiang, Atsumoto Ohashi, Ryuichiro Higashinaka
Full-duplex spoken dialogue models, such as Moshi, enable natural, low-latency voice conversations. However, they remain limited to the audio modality, lacking the facial expressio…