2 citations · 2 across the 13 of their papers we have counts for
11 papers · 1 filter
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs
Zizhong Ding, Junxian Li, Kai Liu +4
Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive e…
TurboClear: One-Step Object-Effect Removal via Region-Calibrated Distribution Matching and Fusion
Jiawei Guo, Junxian Li, Yixin Tang +4
Recently, diffusion-based removal methods have achieved promising visual quality in removing both target objects and their associated effects. However, they typically rely on multi…
Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing
Weiwei Tan, Junxian Li, Rui Wang +3
Unified multimodal models (UMMs) have recently demonstrated powerful instruction-based image editing capabilities, while also raising serious concerns about the unauthorized manipu…
PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models
Yongsen Cheng, Kai Liu, Kaiwen Tao +5
Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challenging in resource-constrained sc…
GTR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models
Junxian Li, Kai Liu, Zizhong Ding +4
The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. In this paper, we focus on un…
FlashClear: Ultra-Fast Image Content Removal via Efficient Step Distillation and Feature Caching
Yixin Tang, Jiawei Guo, Junxian Li +6
Recently, diffusion-based object removal models have achieved impressive results in eliminating objects and their associated visual effects. However, they indiscriminately denoise…