activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study

Weichen Zhang, Ruiying Peng, Xin Zeng +9

3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some promising results, the advantages of p…

cs.CV2025

Progressive Supernet Training for Efficient Visual Autoregressive Modeling

Xiaoyue Chen, Yuling Shi, Kaiyuan Li +5

Visual Auto-Regressive (VAR) models significantly reduce inference steps through the "next-scale" prediction paradigm. However, progressive multi-scale generation incurs substantia…

cs.CV2025

Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization

Kaiyuan Li, Xiaoyue Chen, Chen Gao +2

Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image toke…

cs.CV2025

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent

Jiaao Li, Kaiyuan Li, Chen Gao +2

Egomotion videos are first-person recordings where the view changes continuously due to the agent's movement. As they serve as the primary visual input for embodied AI agents, maki…

cs.CV2025

How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM

Jirong Zha, Yuxuan Fan, Xiao Yang +2

3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs)…

cs.CV2025

PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models

Yu Meng, Kaiyuan Li, Chenran Huang +4

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a range of multimodal tasks. However, their inference efficiency is constrained by the large n…