4 citations · 4 across the 1 of their papers we have counts for
4 papers
ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization
Fanrui Zhang, Jiawei Liu, Jiaying Zhu +4
Multimodal Large Language Models (MLLMs), such as GPT4o, have shown strong capabilities in visual reasoning and explanation generation. However, despite these strengths, they face…
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
Jiaying Zhu, Yurui Zhu, Xin Lu +5
Multimodal Large Language Models (MLLMs) encounter significant computational and memory bottlenecks from the massive number of visual tokens generated by high-resolution images or…
DemosaicFormer: Coarse-to-Fine Demosaicing Network for HybridEVS Camera
Senyan Xu, Zhijing Sun, Jiaying Zhu +3
Hybrid Event-Based Vision Sensor (HybridEVS) is a novel sensor integrating traditional frame-based and event-based sensors, offering substantial benefits for applications requiring…
MIPI 2024 Challenge on Demosaic for HybridEVS Camera: Methods and Results
Yaqi Wu, Zhihao Fan, Xiaofeng Chu +46
The increasing demand for computational photography and imaging on mobile platforms has led to the widespread development and integration of advanced image sensors with novel algor…