14 papers · 1 filter
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs
Zizhong Ding, Junxian Li, Kai Liu +4
Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive e…
Freqformer: Image-Demoiréing Transformer via Effective Frequency Decomposition
Xiaoyang Liu, Bolin Qiu, Zheng Chen +5
Image demoiréing remains a challenging task due to the complex interplay between texture corruption and color distortions caused by moiré patterns. Existing methods, especially t…
PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models
Yongsen Cheng, Kai Liu, Kaiwen Tao +5
Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challenging in resource-constrained sc…
DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models
Xinrui Shi, Kai Liu, Ziqing Zhang +3
Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations…
InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution
Ziqing Zhang, Kai Liu, Zheng Chen +5
Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent challenges when processing long s…
Accelerating Rectified Flow Models via Trajectory-Aware Caching
Xiao Liu, Kai Liu, Naiyang Guan +5
Diffusion and rectified flow (RF) models generate high-fidelity images and videos, but their iterative velocity-field evaluations are computationally expensive. Existing caching me…