most citedGenerate then Select: Open-ended Visual Question Answering Guided by World Knowledge

1 citations · 1 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL2023

Rethinking and Improving Multi-task Learning for End-to-end Speech Translation

Yuhao Zhang, Chen Xu, Bei Li +4

Significant improvements in end-to-end speech translation (ST) have been achieved through the application of multi-task learning. However, the extent to which auxiliary tasks are h…

cs.CL2023

Bridging the Gaps of Both Modality and Language: Synchronous Bilingual CTC for Speech Translation and Speech Recognition

Chen Xu, Xiaoqian Liu, Erfeng He +6

In this study, we present synchronous bilingual Connectionist Temporal Classification (CTC), an innovative framework that leverages dual CTC to bridge the gaps of both modality and…

astro-ph.SR2023

A catalogue and statistical analysis for magnetic stars

Abdurepqet Rustem, Guoliang Lv, Jinzhong Liu +5

Magnetic fields are significant in the structure and evolution of stars. We present a comprehensive catalogue of 1784 known magnetic stars, detailing their identifications, HD numb…

cs.CL20231 cited

Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge

Xingyu Fu, Sheng Zhang, Gukyeong Kwon +10

The open-ended Visual Question Answering (VQA) task requires AI models to jointly reason over visual and natural language inputs using world knowledge. Recently, pre-trained Langua…

cs.CL2023

CTC-based Non-autoregressive Speech Translation

Chen Xu, Xiaoqian Liu, Xiaowen Liu +9

Combining end-to-end speech translation (ST) and non-autoregressive (NAR) generation is promising in language and speech processing for their advantages of less error propagation a…

cs.CL2023

Bridging the Granularity Gap for Acoustic Modeling

Chen Xu, Yuhao Zhang, Chengbo Jiao +7

While Transformer has become the de-facto standard for speech, modeling upon the fine-grained frame-level features remains an open challenge of capturing long-distance dependencies…