12 papers
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Siyu Yan, Zhuoran Yan, Haiying Xu +10
The paper presents See2Think, an evaluation framework and benchmark for testing whether multimodal large language models actually use intermediate visual states during reasoning, a…
StableI2I: Spotting Unintended Changes in Image-to-Image Transition
Jiayang Li, Shuo Cao, Xiaohui Li +6
In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. H…
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
Yi Xin, Siqi Luo, Tianxiang Xu +13
Diffusion Multi-modal Large Language Models (dMLLMs) have recently emerged as a novel architecture unifying image generation and understanding. However, developing effective and ef…
LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution
Xiaohui Li, Shaobin Zhuang, Shuo Cao +6
Generative models for Image Super-Resolution (SR) are increasingly powerful, yet their reliance on self-attention's quadratic complexity (O(N^2)) creates a major computational bott…
Accelerating Masked Image Generation by Learning Latent Controlled Dynamics
Kaiwen Zhu, Quansheng Zeng, Yuandong Pu +8
Masked Image Generation Models (MIGMs) have achieved great success, yet their efficiency is hampered by the multiple steps of bi-directional attention. In fact, there exists notabl…
Toward Generalizable Deblurring: Leveraging Massive Blur Priors with Linear Attention for Real-World Scenarios
Yuanting Gao, Shuo Cao, Xiaohui Li +3
Image deblurring has advanced rapidly with deep learning, yet most methods exhibit poor generalization beyond their training datasets, with performance dropping significantly in re…