From the 1 of 14 linked papers with an AI index.
14 papers
ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models
Xiaolin Chen, Xuemeng Song, Wenhao Shi +3
Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task th…
FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
Bohan Hou, Haoqiang Lin, Xuemeng Song +4
The paper introduces an automated pipeline to create a fine-grained multimodal dataset and a two-stage fine-tuning strategy that improves multimodal large language models' ability…
Diverse-Intent Multi-Turn Fashion Image Retrieval
Mingqiang Tang, Haokun Wen, Meng Liu +3
Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every inte…
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
Wenqi Liu, Yunxiao Wang, Shijie Ma +14
In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To addres…
FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning
Haokun Wen, Xuemeng Song, Xinghao Xie +3
Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is highly desired in practice.…
OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
Yisen Feng, Leigang Qu, Haoyu Zhang +5
In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accu…