2 papers
cs.SD2024
DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
Shengqiang Liu, Da Liu, Anna Wang +3
Analyzing real-world multimodal signals is an essential and challenging task for intelligent voice assistants (IVAs). Mainstream approaches have achieved remarkable performance on…
cs.SD2024
MV: A multi-modal multi-view approach for Device-Directed Speech Detection
Anna Wang, Da Liu, Zhiyu Zhang +3
With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on…