NewEvery arXiv paper, its researchers & institutions — mapped.
information retrieval

VIG-RL: Learning to Search and Insert for Verified Image Grounding

arXiv:2607.28055

summary

The paper introduces VIG-RL, a reinforcement‑learning based agent that dynamically decides when to retrieve, select, and insert authentic images into text responses, improving verified image grounding in knowledge‑intensive tasks.

Abstract

In knowledge-intensive scenarios, providing reliable interleaved text-image responses requires Verified Image Grounding (VIG): the precise integration of retrieved authentic visual evidence. Existing retrieval-augmented frameworks predominantly rely on decoupled, static pipelines, inherently failing to dynamically reason about when external knowledge is required and where visual assets should be contextually inserted. To bridge this gap, we propose VIG-RL, an autonomous agentic framework that formulates the search-selection-insertion workflow as an active decision-making process. Operating within a dynamic ReAct-style loop, VIG-RL is optimized via reinforcement learning, guided by a composite reward system that holistically evaluates the agent's step-by-step tool execution and final multimodal alignment. Extensive evaluations demonstrate that VIG-RL establishes a new state-of-the-art, significantly outperforming existing static baselines.

Topics & keywords

#verified image grounding#reinforcement learning#multimodal retrieval#tool use#dynamic reasoningreinforcement learningimage groundingretrieval‑augmented generationtool selectionmultimodal alignmentreward design