activity
20212024
most citedNICE-Beam: Neural Integrated Covariance Estimators for Time-Varying Beamformers

7 citations · 10 across the 9 of their papers we have counts for

collaborators

9 papers

cs.SD2024

FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses

Zhongweiyang Xu, Ali Aroudi, Ke Tan +4

This paper presents a novel multi-channel speech enhancement approach, FoVNet, that enables highly efficient speech enhancement within a configurable field of view (FoV) of a smart…

cs.SD2024

All Neural Low-latency Directional Speech Extraction

Ashutosh Pandey, Sanha Lee, Juan Azcarreta +2

We introduce a novel all neural model for low-latency directional speech extraction. The model uses direction of arrival (DOA) embeddings from a predefined spatial grid, which are…

eess.AS2024

AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling

Vahid Ahmadi Kalkhorani, Cheng Yu, Anurag Kumar +3

Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an audiovisual (AV) system for speech enhancement, target…

eess.AS2024

A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement

Ravi Shankar, Ke Tan, Buye Xu +1

Self-supervised learned models have been found to be very effective for certain speech tasks such as automatic speech recognition, speaker identification, keyword spotting and othe…

cs.SD2024

On the Importance of Neural Wiener Filter for Resource Efficient Multichannel Speech Enhancement

Tsun-An Hsieh, Jacob Donley, Daniel Wong +2

We introduce a time-domain framework for efficient multichannel speech enhancement, emphasizing low latency and computational efficiency. This framework incorporates two compact de…

cs.SD2024

Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement

Ashutosh Pandey, Buye Xu

We present a novel model designed for resource-efficient multichannel speech enhancement in the time domain, with a focus on low latency, lightweight, and low computational require…