most citedHierarchical Residual Learning Based Vector Quantized Variational Autoencoder for Image Reconstruction and Generation

1 citations · 4 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL20231 cited

Generative error correction for code-switching speech recognition using large language models

Chen Chen, Yuchen Hu, Chao-Han Huck Yang +3

Code-switching (CS) speech refers to the phenomenon of mixing two or more languages within the same sentence. Despite the recent advances in automatic speech recognition (ASR), CS-…

eess.AS2023

Boosting End-to-End Multilingual Phoneme Recognition through Exploiting Universal Speech Attributes Constraints

Hao Yen, Sabato Marco Siniscalchi, Chin-Hui Lee

We propose a first step toward multilingual end-to-end automatic speech recognition (ASR) by integrating knowledge about speech articulators. The key idea is to leverage a rich set…

eess.AS20231 cited

The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction

Shilong Wu, Chenxi Wang, Hang Chen +13

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advan…

cs.MM20231 cited

The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition

Zhe Wang, Shilong Wu, Hang Chen +12

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research…

cs.CV20221 cited

Hierarchical Residual Learning Based Vector Quantized Variational Autoencoder for Image Reconstruction and Generation

Mohammad Adiban, Kalin Stefanov, Sabato Marco Siniscalchi +1

We propose a multi-layer variational autoencoder method, we call HR-VQVAE, that learns hierarchical discrete representations of the data. By utilizing a novel objective function, e…