collaborators

6 papers

cs.CV2025

Progressive Supernet Training for Efficient Visual Autoregressive Modeling

Xiaoyue Chen, Yuling Shi, Kaiyuan Li +5

Visual Auto-Regressive (VAR) models significantly reduce inference steps through the "next-scale" prediction paradigm. However, progressive multi-scale generation incurs substantia…

cs.CV2025

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent

Jiaao Li, Kaiyuan Li, Chen Gao +2

Egomotion videos are first-person recordings where the view changes continuously due to the agent's movement. As they serve as the primary visual input for embodied AI agents, maki…

cs.RO2025

AirScape: An Aerial Generative World Model with Motion Controllability

Baining Zhao, Rongze Tang, Mingyuan Jia +9

How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general s…

cs.CV2025

Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization

Kaiyuan Li, Xiaoyue Chen, Chen Gao +2

Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image toke…

cs.CV2025

PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models

Yu Meng, Kaiyuan Li, Chenran Huang +4

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a range of multimodal tasks. However, their inference efficiency is constrained by the large n…

cs.CV2025

Understanding and Evaluating Hallucinations in 3D Visual Language Models

Ruiying Peng, Kaiyuan Li, Weichen Zhang +3

Recently, 3D-LLMs, which combine point-cloud encoders with large models, have been proposed to tackle complex tasks in embodied intelligence and scene understanding. In addition to…