4 papers
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
Xinyi Chen, Yilun Chen, Yanwei Fu +26
We introduce InternVLA-M1, a unified framework for spatial grounding and robot control that advances instruction-following robots toward scalable, general-purpose intelligence. Its…
VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
Qixin Sun, Ziqin Wang, Hengyuan Zhao +6
Retrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating…
"Hi AirStar, Guide Me to the Badminton Court."
Ziqin Wang, Jinyu Chen, Xiangyi Zheng +3
Unmanned Aerial Vehicles, operating in environments with relatively few obstacles, offer high maneuverability and full three-dimensional mobility. This allows them to rapidly appro…
LLaVA-CMoE: Towards Continual Mixture of Experts for Large Vision-Language Models
Hengyuan Zhao, Ziqin Wang, Qixin Sun +5
Mixture of Experts (MoE) architectures have recently advanced the scalability and adaptability of large language models (LLMs) for continual multimodal learning. However, efficient…