65 citations · 307 across the 32 of their papers we have counts for
4 papers · 2 filters
CASA-ASR: Context-Aware Speaker-Attributed ASR
Mohan Shi, Zhihao Du, Qian Chen +5
Recently, speaker-attributed automatic speech recognition (SA-ASR) has attracted a wide attention, which aims at answering the question ``who spoke what''. Different from modular s…
Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction
Mohan Shi, Yuchun Shu, Lingyun Zuo +4
For speech interaction, voice activity detection (VAD) is often used as a front-end. However, traditional VAD algorithms usually need to wait for a continuous tail silence to reach…
Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection
Xiao-Min Zeng, Yan Song, Zhu Zhuo +5
In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE)…
AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer
Kang Li, Yan Song, Li-Rong Dai +3
In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio t…