2 papers
cs.CV2026
AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs
Aahana Basappa, Pranay Goel, Anusri Karra +3
We investigate visual reasoning limitations of both multimodal large language models (MLLMs) and image generation models (IGMs) by creating a novel benchmark to systematically comp…
cs.CL2025
Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation Generation
Palaash Goel, Dushyant Singh Chauhan, Md Shad Akhtar
Sarcasm is a linguistic phenomenon that intends to ridicule a target (e.g., entity, event, or person) in an inherent way. Multimodal Sarcasm Explanation (MuSE) aims at revealing th…