activity
20172020
most citedConvolutional Gated Recurrent Neural Network Incorporating Spatial Features for Audio Tagging

10 citations · 27 across the 8 of their papers we have counts for

collaborators

12 papers

cs.SD20206 cited

T-vectors: Weakly Supervised Speaker Identification Using Hierarchical Transformer Model

Yanpei Shi, Mingjie Chen, Qiang Huang +1

Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task. This paper proposes a hierarchical network with transformer encoders…

eess.AS2020

Improving Audio Anomalies Recognition Using Temporal Convolutional Attention Network

Qiang Huang, Thomas Hain

Anomalous audio in speech recordings is often caused by speaker voice distortion, external noise, or even electric interferences. These obstacles have become a serious problem in s…

eess.AS2020

Exploration of Audio Quality Assessment and Anomaly Localisation Using Attention Models

Qiang Huang, Thomas Hain

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requir…

eess.AS20201 cited

Speaker Re-identification with Speaker Dependent Speech Enhancement

Yanpei Shi, Qiang Huang, Thomas Hain

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here sp…

eess.AS20201 cited

Weakly Supervised Training of Hierarchical Attention Networks for Speaker Identification

Yanpei Shi, Qiang Huang, Thomas Hain

Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task. In this paper, a hierarchical attention network is proposed to solve…

cs.CL2020

Robust Speaker Recognition Using Speech Enhancement And Attention Model

Yanpei Shi, Qiang Huang, Thomas Hain

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performanc…