35 citations · 94 across the 11 of their papers we have counts for
16 papers
Aligning Source Visual and Target Language Domains for Unpaired Video Captioning
Fenglin Liu, Xian Wu, Chenyu You +3
Training supervised video captioning model requires coupled video-caption pairs. However, for many targeted languages, sufficient paired data are not available. To this end, we int…
Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations
Peng Jin, Jinfa Huang, Fenglin Liu +5
Most video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to t…
DiMBERT: Learning Vision-Language Grounded Representations with Disentangled Multimodal-Attention
Fenglin Liu, Xian Wu, Shen Ge +4
Vision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (…
End-to-end Spoken Conversational Question Answering: Task, Dataset and Model
Chenyu You, Nuo Chen, Fenglin Liu +3
In spoken question answering, the systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that huma…
AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation
Di You, Fenglin Liu, Shen Ge +3
Recently, medical report generation, which aims to automatically generate a long and coherent descriptive paragraph of a given medical image, has received growing research interest…
Audio-Oriented Multimodal Machine Comprehension: Task, Dataset and Model
Zhiqi Huang, Fenglin Liu, Xian Wu +4
While Machine Comprehension (MC) has attracted extensive research interests in recent years, existing approaches mainly belong to the category of Machine Reading Comprehension task…