2 papers
cs.CV2026
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
Jialuo He, Huangxun Chen
Visual token reduction is critical for accelerating Vision-Language Models (VLMs), since visual inputs are represented as token sequences that introduce substantial computational o…
cs.CV2024
ObjVariantEnsemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Objects
Qihang Cao, Huangxun Chen
3D scene understanding is an important task, and there has been a recent surge of research interest in aligning 3D representations of point clouds with text to empower embodied AI.…