35 citations · 95 across the 8 of their papers we have counts for
10 papers
Aligning Source Visual and Target Language Domains for Unpaired Video Captioning
Fenglin Liu, Xian Wu, Chenyu You +3
Training supervised video captioning model requires coupled video-caption pairs. However, for many targeted languages, sufficient paired data are not available. To this end, we int…
Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations
Peng Jin, Jinfa Huang, Fenglin Liu +5
Most video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to t…
DeltaNet:Conditional Medical Report Generation for COVID-19 Diagnosis
Xian Wu, Shuxin Yang, Zhaopeng Qiu +6
Fast screening and diagnosis are critical in COVID-19 patient treatment. In addition to the gold standard RT-PCR, radiological imaging like X-ray and CT also works as an important…
End-to-end Spoken Conversational Question Answering: Task, Dataset and Model
Chenyu You, Nuo Chen, Fenglin Liu +3
In spoken question answering, the systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that huma…
AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation
Di You, Fenglin Liu, Shen Ge +3
Recently, medical report generation, which aims to automatically generate a long and coherent descriptive paragraph of a given medical image, has received growing research interest…
Towards Joint Intent Detection and Slot Filling via Higher-order Attention
Dongsheng Chen, Zhiqi Huang, Xian Wu +2
Intent detection (ID) and Slot filling (SF) are two major tasks in spoken language understanding (SLU). Recently, attention mechanism has been shown to be effective in jointly opti…