4 papers
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
Pehuén Moure, Pehuén Moure, Niclas Pokel +5
Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility of improving performance by co…
Towards Lightweight Adaptation of Speech Enhancement Models in Real-World Environments
Longbiao Cheng, Shih-Chii Liu
Recent studies have shown that post-deployment adaptation can improve the robustness of speech enhancement models in unseen noise conditions. However, existing methods often incur…
Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement
Longbiao Cheng, Ashutosh Pandey, Buye Xu +3
Deep learning-based speech enhancement (SE) methods often face significant computational challenges when needing to meet low-latency requirements because of the increased number of…
DeltaKWS: A 65nm 36nJ/Decision Bio-inspired Temporal-Sparsity-Aware Digital Keyword Spotting IC with 0.6V Near-Threshold SRAM
Qinyu Chen, Kwantae Kim, Chang Gao +4
This paper introduces DeltaKWS, to the best of our knowledge, the first RNN-enabled fine-grained temporal sparsity-aware KWS IC for voice-controlled devices. The 65 nm prototyp…