1 paper
Zehua Cheng, Wei Dai, Jiahao Sun
Multi-modal retrieval-augmented generation (MRAG) systems retrieve visual evidence from large image corpora to ground the responses of large multi-modal models, yet the retrieved i…