6 papers
IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval
Jiale Huang, Zixu Li, Zhiwei Chen +3
Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing methods explore cross-modal cor…
HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval
Zixu Li, Yupeng Hu, Zhiwei Chen +4
Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference im…
Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning
Dongjie Cheng, Yongqi Li, Zhixin Ma +5
Multimodal Large Language Models (MLLMs) are making significant progress in multimodal reasoning. Early approaches focus on pure text-based reasoning. More recent studies have inco…
HINT: Composed Image Retrieval with Dual-path Compositional Contextualized Network
Mingyu Zhang, Zixu Li, Zhiwei Chen +5
Composed Image Retrieval (CIR) is a challenging image retrieval paradigm. It aims to retrieve target images from large-scale image databases that are consistent with the modificati…
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
Chao Chen, Zhixin Ma, Yongqi Li +4
Multimodal reasoning aims to enhance the capabilities of MLLMs by incorporating intermediate reasoning steps before reaching the final answer. It has evolved from text-only reasoni…
MIST: Towards Multi-dimensional Implicit BiaS Evaluation of LLMs for Theory of Mind
Yanlin Li, Hao Liu, Huimin Liu +3
Theory of Mind (ToM) in Large Language Models (LLMs) refers to the model's ability to infer the mental states of others, with failures in this ability often manifesting as systemic…