2 papers
cs.CV2025
MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition
Shu Zhao, Nilesh Ahuja, Tan Yu +2
Pre-trained vision language models have shown remarkable performance on visual recognition tasks, but they typically assume the availability of complete multimodal inputs during bo…
cs.CV2025
Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
Shu Zhao, Tianyi Shen, Nilesh Ahuja +2
Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a promising method to generate factual and up-to-date responses of Multimodal Large Language Models (MLLMs) by incor…