activity
20192023
most citedIntegrating the Data Augmentation Scheme with Various Classifiers for Acoustic Scene Modeling

67 citations · 191 across the 39 of their papers we have counts for

collaborators
Showing 2021Show all

6 papers · 1 filter

cs.SD2021★ 3 cited

Multi-Variant Consistency based Self-supervised Learning for Robust Automatic Speech Recognition

Changfeng Gao, Gaofeng Cheng, Pengyuan Zhang

Automatic speech recognition (ASR) has shown rapid advances in recent years but still degrades significantly in far-field and noisy environments. The recent development of self-sup…

eess.AS2021

Wav2vec-S: Semi-Supervised Pre-Training for Low-Resource ASR

Han Zhu, Li Wang, Jindong Wang +3

Self-supervised pre-training could effectively improve the performance of low-resource automatic speech recognition (ASR). However, existing self-supervised pre-training are task-a…

cs.SD2021★ 1 cited

Adaptive Margin Circle Loss for Speaker Verification

Runqiu Xiao

Deep-Neural-Network (DNN) based speaker verification sys-tems use the angular softmax loss with margin penalties toenhance the intra-class compactness of speaker embeddings,which a…

eess.AS2021★ 8 cited

Improved Conformer-based End-to-End Speech Recognition Using Neural Architecture Search

Yukun Liu, Ta Li, Pengyuan Zhang +1

Recently neural architecture search(NAS) has been successfully used in image classification, natural language processing, and automatic speech recognition(ASR) tasks for finding th…

cs.SD2021

DPT-FSNet: Dual-path Transformer Based Full-band and Sub-band Fusion Network for Speech Enhancement

Feng Dang, Hangting Chen, Pengyuan Zhang

Sub-band models have achieved promising results due to their ability to model local patterns in the spectrogram. Some studies further improve the performance by fusing sub-band and…

eess.AS2021

Beam-Guided TasNet: An Iterative Speech Separation Framework with Multi-Channel Output

Hangting Chen, Yang Yi, Dang Feng +1

Time-domain audio separation network (TasNet) has achieved remarkable performance in blind source separation (BSS). Classic multi-channel speech processing framework employs signal…