6 papers
Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems
Jingjing Jiang, Atsumoto Ohashi, Ryuichiro Higashinaka
Full-duplex spoken dialogue models, such as Moshi, enable natural, low-latency voice conversations. However, they remain limited to the audio modality, lacking the facial expressio…
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
Atsumoto Ohashi, Neil Zeghidour, Alexandre Défossez +1
Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely wi…
Towards a Japanese Full-duplex Spoken Dialogue System
Atsumoto Ohashi, Shinya Iizuka, Jingjing Jiang +1
Full-duplex spoken dialogue systems, which can model simultaneous bidirectional features of human conversations such as speech overlaps and backchannels, have attracted significant…
Universal Post-Processing Networks for Joint Optimization of Modules in Task-Oriented Dialogue Systems
Atsumoto Ohashi, Ryuichiro Higashinaka
Post-processing networks (PPNs) are components that modify the outputs of arbitrary modules in task-oriented dialogue systems and are optimized using reinforcement learning (RL) to…
On the True Distribution Approximation of Minimum Bayes-Risk Decoding
Atsumoto Ohashi, Ukyo Honda, Tetsuro Morimura +1
Minimum Bayes-risk (MBR) decoding has recently gained renewed attention in text generation. MBR decoding considers texts sampled from a model as pseudo-references and selects the t…
JMultiWOZ: A Large-Scale Japanese Multi-Domain Task-Oriented Dialogue Dataset
Atsumoto Ohashi, Ryu Hirai, Shinya Iizuka +1
Dialogue datasets are crucial for deep learning-based task-oriented dialogue system research. While numerous English language multi-domain task-oriented dialogue datasets have been…