2 citations · 2 across the 4 of their papers we have counts for
4 papers
Gated Multimodal Fusion with Contrastive Learning for Turn-taking Prediction in Human-robot Dialogue
Jiudong Yang, Peiying Wang, Yi Zhu +3
Turn-taking, aiming to decide when the next speaker can start talking, is an essential component in building human-robot spoken dialogue systems. Previous studies indicate that mul…
Building Robust Spoken Language Understanding by Cross Attention between Phoneme Sequence and ASR Hypothesis
Zexun Wang, Yuquan Le, Yi Zhu +4
Building Spoken Language Understanding (SLU) robust to Automatic Speech Recognition (ASR) errors is an essential issue for various voice-enabled virtual assistants. Considering tha…
ViDA-MAN: Visual Dialog with Digital Humans
Tong Shen, Jiawei Zuo, Fan Shi +7
We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text o…
Conversational Query Rewriting with Self-supervised Learning
Hang Liu, Meng Chen, Youzheng Wu +2
Context modeling plays a critical role in building multi-turn dialogue systems. Conversational Query Rewriting (CQR) aims to simplify the multi-turn dialogue modeling into a single…