1.5k citations · 1.6k across the 12 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation
Gen Li, Peiyu Liu
Vision-Language Models (VLMs) excel at visual reasoning but still struggle with integrating external knowledge. Retrieval-Augmented Generation (RAG) is a promising solution, but cu…
cs.CV2021
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
Yuqi Huo, Manli Zhang, Guangzhen Liu +32
Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction bet…