23 citations · 26 across the 3 of their papers we have counts for
15 papers
Multi-objective Non-intrusive Hearing-aid Speech Assessment Model
Hsin-Tien Chiang, Szu-Wei Fu, Hsin-Min Wang +2
Without the need for a clean reference, non-intrusive speech assessment methods have caught great attention for objective evaluations. While deep learning models have been used to…
Neural domain alignment for spoken language recognition based on optimal transport
Xugang Lu, Peng Shen, Yu Tsao +1
Domain shift poses a significant challenge in cross-domain spoken language recognition (SLR) by reducing its effectiveness. Unsupervised domain adaptation (UDA) algorithms have bee…
Deep Complex U-Net with Conformer for Audio-Visual Speech Enhancement
Shafique Ahmed, Chia-Wei Chen, Wenze Ren +7
Recent studies have increasingly acknowledged the advantages of incorporating visual data into speech enhancement (SE) systems. In this paper, we introduce a novel audio-visual SE…
The VoiceMOS Challenge 2023: Zero-shot Subjective Speech Quality Prediction for Multiple Domains
Erica Cooper, Wen-Chin Huang, Yu Tsao +3
We present the second edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthesized an…
Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-based ASR
Xugang Lu, Peng Shen, Yu Tsao +1
Due to the modality discrepancy between textual and acoustic modeling, efficiently transferring linguistic knowledge from a pretrained language model (PLM) to acoustic encoding for…
Cross-modal Alignment with Optimal Transport for CTC-based ASR
Xugang Lu, Peng Shen, Yu Tsao +1
Temporal connectionist temporal classification (CTC)-based automatic speech recognition (ASR) is one of the most successful end to end (E2E) ASR frameworks. However, due to the tok…