most citedAdaptive Few-Shot Learning Algorithm for Rare Sound Event Detection

4 citations · 5 across the 7 of their papers we have counts for

collaborators

7 papers

cs.SD2022

Learning Invariant Representation and Risk Minimized for Unsupervised Accent Domain Adaptation

Chendong Zhao, Jianzong Wang, Xiaoyang Qu +2

Unsupervised representation learning for speech audios attained impressive performances for speech recognition tasks, particularly when annotated speech is limited. However, the un…

cs.CL2022

Adaptive Sparse and Monotonic Attention for Transformer-based Automatic Speech Recognition

Chendong Zhao, Jianzong Wang, Wen qi Wei +3

The Transformer architecture model, based on self-attention and multi-head attention, has achieved remarkable success in offline end-to-end Automatic Speech Recognition (ASR). Howe…

cs.SD2022

DT-SV: A Transformer-based Time-domain Approach for Speaker Verification

Nan Zhang, Jianzong Wang, Zhenhou Hong +3

Speaker verification (SV) aims to determine whether the speaker's identity of a test utterance is the same as the reference speech. In the past few years, extracting speaker embedd…

quant-ph2022

QSpeech: Low-Qubit Quantum Speech Application Toolkit

Zhenhou Hong, Jianzong Wang, Xiaoyang Qu +3

Quantum devices with low qubits are common in the Noisy Intermediate-Scale Quantum (NISQ) era. However, Quantum Neural Network (QNN) running on low-qubit quantum devices would be d…

cs.SD20224 cited

Adaptive Few-Shot Learning Algorithm for Rare Sound Event Detection

Chendong Zhao, Jianzong Wang, Leilai Li +2

Sound event detection is to infer the event by understanding the surrounding environmental sounds. Due to the scarcity of rare sound events, it becomes challenging for the well-tra…

eess.AS2022

r-G2P: Evaluating and Enhancing Robustness of Grapheme to Phoneme Conversion by Controlled noise introducing and Contextual information incorporation

Chendong Zhao, Jianzong Wang, Xiaoyang Qu +2

Grapheme-to-phoneme (G2P) conversion is the process of converting the written form of words to their pronunciations. It has an important role for text-to-speech (TTS) synthesis and…