activity
20152022
most citedWaveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension

65 citations · 294 across the 25 of their papers we have counts for

collaborators

26 papers

cs.CL20211 cited

The USTC-NELSLIP Systems for Simultaneous Speech Translation Task at IWSLT 2021

Dan Liu, Mengge Du, Xiaoxi Li +2

This paper describes USTC-NELSLIP's submissions to the IWSLT2021 Simultaneous Speech Translation task. We proposed a novel simultaneous translation model, Cross Attention Augmented…

eess.AS20218 cited

XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition

Zi-Qiang Zhang, Yan Song, Ming-Hui Wu +2

In this paper, we propose a weakly supervised multilingual representation learning framework, called cross-lingual self-training (XLST). XLST is able to utilize a small amount of a…

cs.CV20204 cited

Lip-reading with Hierarchical Pyramidal Convolution and Self-Attention

Hang Chen, Jun Du, Yu Hu +3

In this paper, we propose a novel deep learning architecture to improving word-level lip-reading. On the one hand, we first introduce the multi-scale processing into the spatial fe…

cs.SD2020

Correlating Subword Articulation with Lip Shapes for Embedding Aware Audio-Visual Speech Enhancement

Hang Chen, Jun Du, Yu Hu +3

In this paper, we propose a visual embedding approach to improving embedding aware speech enhancement (EASE) by synchronizing visual lip frames at the phone and place of articulati…

eess.AS20206 cited

Voice Conversion by Cascading Automatic Speech Recognition and Text-to-Speech Synthesis with Prosody Transfer

Jing-Xuan Zhang, Li-Juan Liu, Yan-Nian Chen +4

With the development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) technique, it's intuitive to construct a voice conversion system by cascading an ASR a…

eess.AS20209 cited

Attentive Fusion Enhanced Audio-Visual Encoding for Transformer Based Robust Speech Recognition

Liangfa Wei, Jie Zhang, Junfeng Hou +1

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore…