4 papers
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
Haolin Deng, Xin Zou, Zhiwei Jin +3
Multimodal hallucination remains a persistent challenge for Vision-Language Models (VLMs). Standard textual Direct Preference Optimization (DPO) often fails to mitigate it due to a…
PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
Nan Wang, Zhiwei Jin, Chen Chen +1
Document understanding and GUI interaction are among the highest-value applications of Vision-Language Models (VLMs), yet they impose exceptionally heavy computational burden: fine…
Accelerating Learned Video Compression via Low-Resolution Representation Learning
Zidian Qiu, Zongyao He, Zhi Jin
In recent years, the field of learned video compression has witnessed rapid advancement, exemplified by the latest neural video codecs DCVC-DC that has outperformed the upcoming ne…
Latent Modulated Function for Computational Optimal Continuous Image Representation
Zongyao He, Zhi Jin
The recent work Local Implicit Image Function (LIIF) and subsequent Implicit Neural Representation (INR) based works have achieved remarkable success in Arbitrary-Scale Super-Resol…