4 papers
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
Jiaying Zhu, Yurui Zhu, Xin Lu +5
Multimodal Large Language Models (MLLMs) encounter significant computational and memory bottlenecks from the massive number of visual tokens generated by high-resolution images or…
NTIRE 2025 Image Shadow Removal Challenge Report
Florin-Alexandru Vasluianu, Tim Seizinger, Zhuyun Zhou +79
This work examines the findings of the NTIRE 2025 Shadow Removal Challenge. A total of 306 participants have registered, with 17 teams successfully submitting their solutions durin…
ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization
Fanrui Zhang, Jiawei Liu, Jiaying Zhu +4
Multimodal Large Language Models (MLLMs), such as GPT4o, have shown strong capabilities in visual reasoning and explanation generation. However, despite these strengths, they face…
FourierMamba: Fourier Learning Integration with State Space Models for Image Deraining
Dong Li, Yidi Liu, Xueyang Fu +2
Image deraining aims to remove rain streaks from rainy images and restore clear backgrounds. Currently, some research that employs the Fourier transform has proved to be effective…