From the 1 of 5 linked papers with an AI index.
5 papers
FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
Bohan Hou, Haoqiang Lin, Xuemeng Song +4
The paper introduces an automated pipeline to create a fine-grained multimodal dataset and a two-stage fine-tuning strategy that improves multimodal large language models' ability…
A Survey on Video Temporal Grounding with Multimodal Large Language Model
Jianlong Wu, Wei Liu, Ye Liu +4
The recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs).…
Object-Shot Enhanced Grounding Network for Egocentric Video
Yisen Feng, Haoyu Zhang, Meng Liu +2
Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the dis…
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
Haoyu Zhang, Meng Liu, Yisen Feng +3
In contrast to conventional visual question answering, video-grounded dialog necessitates a profound understanding of both dialog history and video content for accurate response ge…
Sentiment-enhanced Graph-based Sarcasm Explanation in Dialogue
Kun Ouyang, Liqiang Jing, Xuemeng Song +3
Sarcasm Explanation in Dialogue (SED) is a new yet challenging task, which aims to generate a natural language explanation for the given sarcastic dialogue that involves multiple m…