activity
20242026
collaborators

7 papers

cs.RO2026

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

Ziqin Wang, Hao Li, Weijun Wang +5

Existing robot datasets remain expensive to curate, embodiment-specific, and insufficiently annotated with the fine-grained structure required for generalizable reasoning, executio…

cs.RO2026

RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation

Hao Li, Ziqin Wang, Zi-han Ding +9

Advances in large vision-language models (VLMs) have stimulated growing interest in vision-language-action (VLA) systems for robot manipulation. However, existing manipulation data…

cs.RO2025

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Xinyi Chen, Yilun Chen, Yanwei Fu +26

We introduce InternVLA-M1, a unified framework for spatial grounding and robot control that advances instruction-following robots toward scalable, general-purpose intelligence. Its…

cs.CL2025

VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples

Qixin Sun, Ziqin Wang, Hengyuan Zhao +6

Retrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating…

cs.RO2025

"Hi AirStar, Guide Me to the Badminton Court."

Ziqin Wang, Jinyu Chen, Xiangyi Zheng +3

Unmanned Aerial Vehicles, operating in environments with relatively few obstacles, offer high maneuverability and full three-dimensional mobility. This allows them to rapidly appro…

cs.CL2025

LLaVA-CMoE: Towards Continual Mixture of Experts for Large Vision-Language Models

Hengyuan Zhao, Ziqin Wang, Qixin Sun +5

Mixture of Experts (MoE) architectures have recently advanced the scalability and adaptability of large language models (LLMs) for continual multimodal learning. However, efficient…