5 papers
Micro Language Models Enable Instant Responses
Wen Cheng, Tuochao Chen, Karim Helwani +3
Edge devices such as smartwatches and smart glasses cannot continuously run even the smallest 100M-1B parameter language models due to power and compute constraints, yet cloud infe…
A Hierarchical End-of-Turn Model with Primary Speaker Segmentation for Real-Time Conversational AI
Karim Helwani, Hoang Do, James Luan +1
We present a real-time front-end for voice-based conversational AI to enable natural turn-taking in two-speaker scenarios by combining primary speaker segmentation with hierarchica…
Sound Source Separation Using Latent Variational Block-Wise Disentanglement
Karim Helwani, Masahito Togami, Paris Smaragdis +1
While neural network approaches have made significant strides in resolving classical signal processing problems, it is often the case that hybrid approaches that draw insight from…
Real-time Stereo Speech Enhancement with Spatial-Cue Preservation based on Dual-Path Structure
Masahito Togami, Jean-Marc Valin, Karim Helwani +3
We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source…
Neural Harmonium: An Interpretable Deep Structure for Nonlinear Dynamic System Identification with Application to Audio Processing
Karim Helwani, Erfan Soltanmohammadi, Michael M. Goodwin
Improving the interpretability of deep neural networks has recently gained increased attention, especially when the power of deep learning is leveraged to solve problems in physics…