4 papers
Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement
Xulin Fan, Juan Azcarreta, Ashutosh Pandey +5
Low-latency, low-compute speech enhancement is essential for wearable devices with real-time communication requirements, but strict computational constraints significantly limit on…
OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics
Wei Chu, Yuanzhe Dong, Ke Tan +7
OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podca…
Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
Heitor R. Guimarães, Ke Tan, Juan Azcarreta +4
Deploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learnin…
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
Joanna Hong, Sanjeel Parekh, Honglie Chen +4
Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in perfor…