2 papers
cs.CV2026
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
Vishal Narnaware, Animesh Gupta, Kevin Zhai +2
Multimodal Diffusion Large Language Models (MDLLMs) achieve high-concurrency generation through parallel masked decoding, yet the architectures remain prone to multimodal hallucina…
cs.CV2026
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
Vishal Narnaware, Ashmal Vayani, Rohit Gupta +2
Stereotype biases in Large Multimodal Models (LMMs) perpetuate harmful societal prejudices, undermining the fairness and equity of AI applications. As LMMs grow increasingly influe…