Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
Chika Maduabuchi, Hao Chen, Yujin Han +1
Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where…
cs.CV2025
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
Jiale Li, Mingrui Wu, Zixiang Jin +5
Despite growing interest in hallucination in Multimodal Large Language Models, existing studies primarily focus on single-image settings, leaving hallucination in multi-image scena…
cs.CV2025
On the robustness of multimodal language model towards distractions
Ming Liu, Hao Chen, Jindong Wang +1
Although vision-language models (VLMs) have achieved significant success in various applications such as visual question answering, their resilience to prompt variations remains an…