activity
20172022
most citedOn Mean Absolute Error for Deep Neural Network Based Vector-to-Vector Regression

303 citations · 631 across the 26 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD2022

A Study of Designing Compact Audio-Visual Wake Word Spotting System Based on Iterative Fine-Tuning in Neural Network Pruning

Hengshun Zhou, Jun Du, Chao-Han Huck Yang +2

Audio-only-based wake word spotting (WWS) is challenging under noisy conditions due to environmental interference in signal transmission. In this paper, we investigate on designing…

cs.SD2021

AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Yihui Fu, Luyao Cheng, Shubo Lv +10

In this paper, we present AISHELL-4, a sizable real-recorded Mandarin speech dataset collected by 8-channel circular microphone array for speech processing in conference scenario.…

cs.SD202120 cited

USTC-NELSLIP System Description for DIHARD-III Challenge

Yuxuan Wang, Maokui He, Shutong Niu +6

This system description describes our submission system to the Third DIHARD Speech Diarization Challenge. Besides the traditional clustering based system, the innovation of our sys…

cs.SD20203 cited

Frequency Gating: Improved Convolutional Neural Networks for Speech Enhancement in the Time-Frequency Domain

Koen Oostermeijer, Qing Wang, Jun Du

One of the strengths of traditional convolutional neural networks (CNNs) is their inherent translational invariance. However, for the task of speech enhancement in the time-frequen…

cs.SD2020

A Two-Stage Approach to Device-Robust Acoustic Scene Classification

Hu Hu, Chao-Han Huck Yang, Xianjun Xia +13

To improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convol…

cs.SD2020

Correlating Subword Articulation with Lip Shapes for Embedding Aware Audio-Visual Speech Enhancement

Hang Chen, Jun Du, Yu Hu +3

In this paper, we propose a visual embedding approach to improving embedding aware speech enhancement (EASE) by synchronizing visual lip frames at the phone and place of articulati…