26 papers
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
Renjie Lu, Xulong Zhang, Xiaoyang Qu +2
Unified Multimodal models (UMMs) built on a single architecture have shown impressive performance in both understanding and generation. We identify a fundamental challenge that lie…
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
Wei Tao, Xiaoyang Qu, Peiqiang Wang +4
Recently, video language models (VLMs) have been applied in various fields. However, the visual token sequence of the VLM is too long, which may cause intolerant inference latency…
Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization
Jianzong Wang, Botao Zhao, Yayun He +2
Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face limitations such as extensive tr…
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
Jiaqi Shi, Yuechan Li, Xulong Zhang +2
High-resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Existing acceleration strategi…
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
Chuhang Liu, Yayun He, Zuheng Kang +2
Vision-Language-Action (VLA) models integrate visual perception, language understanding, and action decision-making for cross-modal semantic alignment, exhibiting broad application…
Confusion-Aware In-Context-Learning for Vision-Language Models in Robotic Manipulation
Yayun He, Zuheng Kang, Botao Zhao +3
Vision-language models (VLMs) have significantly improved the generalization capabilities of robotic manipulation. However, VLM-based systems often suffer from a lack of robustness…