5 papers
RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference
Ben Wan, Yan Feng, Zihan Tang +4
DeepSeek-OCR leverages visual-text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundant textual and structural inf…
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
Zitong Xu, Huiyu Duan, Bingnan Liu +9
The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, edit…
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
Xiaorong Zhu, Ziheng Jia, Jiarui Wang +6
The rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assess…
Exploring bidirectional bounds for minimax-training of Energy-based models
Cong Geng, Jia Wang, Li Chen +3
Energy-based models (EBMs) estimate unnormalized densities in an elegant framework, but they are generally difficult to train. Recent work has linked EBMs to generative adversarial…
Pruning for Sparse Diffusion Models based on Gradient Flow
Ben Wan, Tianyi Zheng, Zhaoyu Chen +2
Diffusion Models (DMs) have impressive capabilities among generation models, but are limited to slower inference speeds and higher computational costs. Previous works utilize one-s…