activity
20182022
most citedTMS: A Temporal Multi-scale Backbone Design for Speaker Embedding

5 citations · 12 across the 8 of their papers we have counts for

collaborators

15 papers

eess.AS20222 cited

Partial Coupling of Optimal Transport for Spoken Language Identification

Xugang Lu, Peng Shen, Yu Tsao +1

In order to reduce domain discrepancy to improve the performance of cross-domain spoken language identification (SLID) system, as an unsupervised domain adaptation (UDA) method, we…

cs.SD20225 cited

TMS: A Temporal Multi-scale Backbone Design for Speaker Embedding

Ruiteng Zhang, Jianguo Wei, Xugang Lu +6

Speaker embedding is an important front-end module to explore discriminative speaker features for many speech applications where speaker information is needed. Current SOTA backbon…

eess.AS2022

A Novel Temporal Attentive-Pooling based Convolutional Recurrent Architecture for Acoustic Signal Enhancement

Tassadaq Hussain, Wei-Chien Wang, Mandar Gogate +5

In acoustic signal processing, the target signals usually carry semantic information, which is encoded in a hierarchal structure of short and long-term contexts. However, the backg…

eess.AS20212 cited

Siamese Neural Network with Joint Bayesian Model Structure for Speaker Verification

Xugang Lu, Peng Shen, Yu Tsao +1

Generative probability models are widely used for speaker verification (SV). However, the generative models are lack of discriminative feature selection ability. As a hypothesis te…

cs.SD2021

MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement

Szu-Wei Fu, Cheng Yu, Tsun-An Hsieh +4

The discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory. Ob…

eess.AS2021

EMA2S: An End-to-End Multimodal Articulatory-to-Speech System

Yu-Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang +4

Synthesized speech from articulatory movements can have real-world use for patients with vocal cord disorders, situations requiring silent speech, or in high-noise environments. In…