3 papers
cs.CL2022
Improving End-to-End Text Image Translation From the Auxiliary Text Translation Task
Cong Ma, Yaping Zhang, Mei Tu +4
End-to-end text image translation (TIT), which aims at translating the source language embedded in images to the target language, has attracted intensive attention in recent resear…
eess.AS2018
Deep Segment Attentive Embedding for Duration Robust Speaker Verification
Bin Liu, Shuai Nie, Yaping Zhang +2
LSTM-based speaker verification usually uses a fixed-length local segment randomly truncated from an utterance to learn the utterance-level speaker embedding, while using the avera…
cs.SD2018
Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training
Bin Liu, Shuai Nie, Yaping Zhang +3
In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) system…