collaborators

5 papers

cs.LG2026

Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach

Yaoxin Yang, Peng Ye, Xudong Tan +4

Multimodal large language models suffer from substantial inference overhead since multimodal KV Cache grows proportionally with the visual input length. Existing multimodal KV Cach…

cs.CV2025

Sequential Token Merging: Revisiting Hidden States

Yan Wen, Peng Ye, Lin Zhang +4

Vision Mambas (ViMs) achieve remarkable success with sub-quadratic complexity, but their efficiency remains constrained by quadratic token scaling with image resolution. While exis…

cs.CV2025

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Xudong Tan, Yaoxin Yang, Peng Ye +5

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for general-purpose robot control through natural language instructions. However, their high inference cost-…

cs.CV2025

TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models

Xudong Tan, Peng Ye, Chongjun Tu +5

Multimodal Large Language Models (MLLMs) are becoming increasingly popular, while the high computational cost associated with multimodal data input, particularly from visual tokens…

cs.CV2025

Multi-Level Decoupled Relational Distillation for Heterogeneous Architectures

Yaoxin Yang, Peng Ye, Weihao Lin +4

Heterogeneous distillation is an effective way to transfer knowledge from cross-architecture teacher models to student models. However, existing heterogeneous distillation methods…