3 papers
cs.CV2026
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs
Zizhong Ding, Junxian Li, Kai Liu +4
Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive e…
cs.CV2026
GTR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models
Junxian Li, Kai Liu, Zizhong Ding +4
The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. In this paper, we focus on un…
cs.CV2026
Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning
Hao Sun, Zi-Jun Ding, Da-Wei Zhou
Class-Incremental Learning (CIL) enables models to continuously integrate new knowledge while mitigating catastrophic forgetting. Driven by the remarkable generalization of CLIP, l…