1 paper · 1 filter
Xinxi Chen, Tianyang Chen, Lijia Hong
We propose a method to improve Visual Question Answering (VQA) with Retrieval-Augmented Generation (RAG) by introducing text-grounded object localization. Rather than retrieving in…