3 papers
cs.CV2026
Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs
Song Zhang, Yanlong Chen, Yilin Li +4
Remote sensing vision-language models (RS-VLMs) face a fundamental mismatch with natural-image counterparts: the same geographic object exhibits radically different visual evidence…
cs.CL2025
VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
Qixin Sun, Ziqin Wang, Hengyuan Zhao +6
Retrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating…
cs.CL2025
LLaVA-CMoE: Towards Continual Mixture of Experts for Large Vision-Language Models
Hengyuan Zhao, Ziqin Wang, Qixin Sun +5
Mixture of Experts (MoE) architectures have recently advanced the scalability and adaptability of large language models (LLMs) for continual multimodal learning. However, efficient…