3 papers
cs.SD2024
DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
Shengqiang Liu, Da Liu, Anna Wang +3
Analyzing real-world multimodal signals is an essential and challenging task for intelligent voice assistants (IVAs). Mainstream approaches have achieved remarkable performance on…
cs.SD2024
MV: A multi-modal multi-view approach for Device-Directed Speech Detection
Anna Wang, Da Liu, Zhiyu Zhang +3
With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on…
cs.LG2024
Turbo your multi-modal classification with contrastive learning
Zhiyu Zhang, Da Liu, Shengqiang Liu +3
Contrastive learning has become one of the most impressive approaches for multi-modal representation learning. However, previous multi-modal works mainly focused on cross-modal und…