5 citations · 12 across the 8 of their papers we have counts for
15 papers
Partial Coupling of Optimal Transport for Spoken Language Identification
Xugang Lu, Peng Shen, Yu Tsao +1
In order to reduce domain discrepancy to improve the performance of cross-domain spoken language identification (SLID) system, as an unsupervised domain adaptation (UDA) method, we…
TMS: A Temporal Multi-scale Backbone Design for Speaker Embedding
Ruiteng Zhang, Jianguo Wei, Xugang Lu +6
Speaker embedding is an important front-end module to explore discriminative speaker features for many speech applications where speaker information is needed. Current SOTA backbon…
A Novel Temporal Attentive-Pooling based Convolutional Recurrent Architecture for Acoustic Signal Enhancement
Tassadaq Hussain, Wei-Chien Wang, Mandar Gogate +5
In acoustic signal processing, the target signals usually carry semantic information, which is encoded in a hierarchal structure of short and long-term contexts. However, the backg…
Siamese Neural Network with Joint Bayesian Model Structure for Speaker Verification
Xugang Lu, Peng Shen, Yu Tsao +1
Generative probability models are widely used for speaker verification (SV). However, the generative models are lack of discriminative feature selection ability. As a hypothesis te…
MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement
Szu-Wei Fu, Cheng Yu, Tsun-An Hsieh +4
The discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory. Ob…
EMA2S: An End-to-End Multimodal Articulatory-to-Speech System
Yu-Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang +4
Synthesized speech from articulatory movements can have real-world use for patients with vocal cord disorders, situations requiring silent speech, or in high-noise environments. In…