14 papers · 1 filter
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs
Zizhong Ding, Junxian Li, Kai Liu +4
Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive e…
PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models
Yongsen Cheng, Kai Liu, Kaiwen Tao +5
Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challenging in resource-constrained sc…
DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models
Xinrui Shi, Kai Liu, Ziqing Zhang +3
Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations…
Accelerating Rectified Flow Models via Trajectory-Aware Caching
Xiao Liu, Kai Liu, Naiyang Guan +5
Diffusion and rectified flow (RF) models generate high-fidelity images and videos, but their iterative velocity-field evaluations are computationally expensive. Existing caching me…
GTR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models
Junxian Li, Kai Liu, Zizhong Ding +4
The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. In this paper, we focus on un…
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
Junxian Li, Kai Liu, Leyang Chen +7
Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting compute…