3 papers
cs.CV2026
MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models
Yuansheng Gao, Wenbin Xing, Jiahao Yuan +4
Video Large Multimodal Models have achieved remarkable progress in video understanding, yet they remain prone to hallucinations, where generated responses are not faithfully suppor…
cs.CV2026
A Unified Perspective on Adversarial Membership Manipulation in Vision Models
Ruize Gao, Kaiwen Zhou, Yongqiang Chen +1
Membership inference attacks (MIAs) aim to determine whether a specific data point was part of a model's training set, serving as effective tools for evaluating privacy leakage of…
cs.CV2024
Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration
Haoze Sun, Wenbo Li, Jiayue Liu +7
Generalization has long been a central challenge in real-world image restoration. While recent diffusion-based restoration methods, which leverage generative priors from text-to-im…