#multimodal retrieval
7 papers · 1 filter
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Yao Xiao, Reuben Tan, Zhen Zhu +3
ReToken introduces a single learnable embedding that acts as a retrieval token to select a sparse set of relevant visual tokens from a cached representation, improving vision-langu…
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
Jiacheng Tao, Qingyun Sun, Haonan Yuan +2
The paper introduces DualG-MRAG, a framework that separates global reasoning and fine-grained evidence matching using macro and micro graphs to improve multimodal retrieval-augment…
VIG-RL: Learning to Search and Insert for Verified Image Grounding
Qinhan Yu, Jun Guang, Chong Chen +1
The paper introduces VIG-RL, a reinforcement‑learning based agent that dynamically decides when to retrieve, select, and insert authentic images into text responses, improving veri…
Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
Xinyu Luo, Hui Liu, Yihua Shao +3
The paper introduces Conditional Retrieval Alignment (CoRA), a gradient‑free method that turns a frozen encoder into a task‑conditioned retriever for on‑device in‑context learning,…
Fine-Grained Food Image Understanding via Target-Aware Data Alignment
Jui-Feng Chi, Wei-Lun Chu, Bruce Coburn +2
The paper introduces a data-centric approach that selects and refines web-collected image‑caption pairs to better train CLIP‑style vision‑language models for fine‑grained food reco…
MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation
Yi Lin, Yihao Ding, Elana Benishay +8
MonteRET is an AI system that combines whole‑volume CT features with region‑level anatomical information and knowledge retrieval to automatically generate more complete and clinica…