297 citations · 623 across the 48 of their papers we have counts for
20 papers
LION : Empowering Multimodal Large Language Model with Dual-Level Visual Knowledge
Gongwei Chen, Leyang Shen, Rui Shao +2
Multimodal Large Language Models (MLLMs) have endowed LLMs with the ability to perceive and understand multi-modal signals. However, most of the existing MLLMs mainly adopt vision…
An Empirical Study of Frame Selection for Text-to-Video Retrieval
Mengxia Wu, Min Cao, Yang Bai +4
Text-to-video retrieval (TVR) aims to find the most relevant video in a large video gallery given a query text. The intricate and abundant context of the video challenges the perfo…
Dual Semantic Knowledge Composed Multimodal Dialog Systems
Xiaolin Chen, Xuemeng Song, Yinwei Wei +2
Textual response generation is an essential task for multimodal task-oriented dialog systems.Although existing studies have achieved fruitful progress, they still suffer from two c…
Stylized Data-to-Text Generation: A Case Study in the E-Commerce Domain
Liqiang Jing, Xuemeng Song, Xuming Lin +3
Existing data-to-text generation efforts mainly focus on generating a coherent text from non-linguistic input data, such as tables and attribute-value pairs, but overlook that diff…
OFAR: A Multimodal Evidence Retrieval Framework for Illegal Live-streaming Identification
Lin Dengtian, Ma Yang, Li Yuhong +3
Illegal live-streaming identification, which aims to help live-streaming platforms immediately recognize the illegal behaviors in the live-streaming, such as selling precious and e…
Learnable Pillar-based Re-ranking for Image-Text Retrieval
Leigang Qu, Meng Liu, Wenjie Wang +3
Image-text retrieval aims to bridge the modality gap and retrieve cross-modal content based on semantic similarities. Prior work usually focuses on the pairwise relations (i.e., wh…