#multimodal retrieval

topicmultimodal retrieval

7 papers · 1 filter

cs.CV2026

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Yao Xiao, Reuben Tan, Zhen Zhu +3

ReToken introduces a single learnable embedding that acts as a retrieval token to select a sparse set of relevant visual tokens from a cached representation, improving vision-langu…

cs.AI2026

DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation

Jiacheng Tao, Qingyun Sun, Haonan Yuan +2

The paper introduces DualG-MRAG, a framework that separates global reasoning and fine-grained evidence matching using macro and micro graphs to improve multimodal retrieval-augment…

cs.IR2026

VIG-RL: Learning to Search and Insert for Verified Image Grounding

Qinhan Yu, Jun Guang, Chong Chen +1

The paper introduces VIG-RL, a reinforcement‑learning based agent that dynamically decides when to retrieve, select, and insert authentic images into text responses, improving veri…

cs.CL2026

Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning

Xinyu Luo, Hui Liu, Yihua Shao +3

The paper introduces Conditional Retrieval Alignment (CoRA), a gradient‑free method that turns a frozen encoder into a task‑conditioned retriever for on‑device in‑context learning,…

cs.CV2026

Fine-Grained Food Image Understanding via Target-Aware Data Alignment

Jui-Feng Chi, Wei-Lun Chu, Bruce Coburn +2

The paper introduces a data-centric approach that selects and refines web-collected image‑caption pairs to better train CLIP‑style vision‑language models for fine‑grained food reco…

cs.CV2026

MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation

Yi Lin, Yihao Ding, Elana Benishay +8

MonteRET is an AI system that combines whole‑volume CT features with region‑level anatomical information and knowledge retrieval to automatically generate more complete and clinica…