4 papers
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
Jiaying Zhu, Yurui Zhu, Xin Lu +5
Multimodal Large Language Models (MLLMs) encounter significant computational and memory bottlenecks from the massive number of visual tokens generated by high-resolution images or…
Elucidating and Endowing the Diffusion Training Paradigm for General Image Restoration
Xin Lu, Xueyang Fu, Jie Xiao +3
While diffusion models demonstrate strong generative capabilities in image restoration (IR) tasks, their complex architectures and iterative processes limit their practical applica…
DemosaicFormer: Coarse-to-Fine Demosaicing Network for HybridEVS Camera
Senyan Xu, Zhijing Sun, Jiaying Zhu +3
Hybrid Event-Based Vision Sensor (HybridEVS) is a novel sensor integrating traditional frame-based and event-based sensors, offering substantial benefits for applications requiring…
MIPI 2024 Challenge on Demosaic for HybridEVS Camera: Methods and Results
Yaqi Wu, Zhihao Fan, Xiaofeng Chu +46
The increasing demand for computational photography and imaging on mobile platforms has led to the widespread development and integration of advanced image sensors with novel algor…