activity
20202025
most citedAdaptive Few-Shot Learning Algorithm for Rare Sound Event Detection

4 citations · 12 across the 16 of their papers we have counts for

collaborators

14 papers

cs.SD2022

Learning Invariant Representation and Risk Minimized for Unsupervised Accent Domain Adaptation

Chendong Zhao, Jianzong Wang, Xiaoyang Qu +2

Unsupervised representation learning for speech audios attained impressive performances for speech recognition tasks, particularly when annotated speech is limited. However, the un…

cs.CV20221 cited

Pose Guided Human Image Synthesis with Partially Decoupled GAN

Jianhan Wu, Jianzong Wang, Shijing Si +2

Pose Guided Human Image Synthesis (PGHIS) is a challenging task of transforming a human image from the reference pose to a target pose while preserving its style. Most existing met…

cs.CL2022

Adaptive Sparse and Monotonic Attention for Transformer-based Automatic Speech Recognition

Chendong Zhao, Jianzong Wang, Wen qi Wei +3

The Transformer architecture model, based on self-attention and multi-head attention, has achieved remarkable success in offline end-to-end Automatic Speech Recognition (ASR). Howe…

cs.CL2022

Blur the Linguistic Boundary: Interpreting Chinese Buddhist Sutra in English via Neural Machine Translation

Denghao Li, Yuqiao Zeng, Jianzong Wang +5

Buddhism is an influential religion with a long-standing history and profound philosophy. Nowadays, more and more people worldwide aspire to learn the essence of Buddhism, attachin…

eess.AS2022

Boosting Star-GANs for Voice Conversion with Contrastive Discriminator

Shijing Si, Jianzong Wang, Xulong Zhang +3

Nonparallel multi-domain voice conversion methods such as the StarGAN-VCs have been widely applied in many scenarios. However, the training of these models usually poses a challeng…

cs.SD2022

DT-SV: A Transformer-based Time-domain Approach for Speaker Verification

Nan Zhang, Jianzong Wang, Zhenhou Hong +3

Speaker verification (SV) aims to determine whether the speaker's identity of a test utterance is the same as the reference speech. In the past few years, extracting speaker embedd…