activity
20182023
most citedA Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition

16 citations · 16 across the 2 of their papers we have counts for

collaborators

8 papers

cs.CL2023

A Knowledge-enhanced Two-stage Generative Framework for Medical Dialogue Information Extraction

Zefa Hu, Ziyi Ni, Jing Shi +2

This paper focuses on term-status pair extraction from medical dialogues (MD-TSPE), which is essential in diagnosis dialogue systems and the automatic scribe of electronic medical…

cs.CV2022

Improving Cross-Modal Understanding in Visual Dialog via Contrastive Learning

Feilong Chen, Xiuyi Chen, Shuang Xu +1

Visual Dialog is a challenging vision-language task since the visual dialog agent needs to answer a series of questions after reasoning over both the image content and dialog histo…

cs.CL2020

"Listen, Understand and Translate": Triple Supervision Decouples End-to-end Speech-to-text Translation

Qianqian Dong, Rong Ye, Mingxuan Wang +4

An end-to-end speech-to-text translation (ST) takes audio in a source language and outputs the text in a target language. Existing methods are limited by the amount of parallel cor…

eess.AS202016 cited

A Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition

Linhao Dong, Cheng Yi, Jianzong Wang +4

End-to-end models are gaining wider attention in the field of automatic speech recognition (ASR). One of their advantages is the simplicity of building that directly recognizes the…

cs.SD2018

Single-channel Speech Dereverberation via Generative Adversarial Training

Chenxing Li, Tieqiang Wang, Shuang Xu +1

In this paper, we propose a single-channel speech dereverberation system (DeReGAT) based on convolutional, bidirectional long short-term memory and deep feed-forward neural network…

eess.AS2018

Multilingual End-to-End Speech Recognition with A Single Transformer on Low-Resource Languages

Shiyu Zhou, Shuang Xu, Bo Xu

Sequence-to-sequence attention-based models integrate an acoustic, pronunciation and language model into a single neural network, which make them very suitable for multilingual aut…