3 papers
cs.LG2026
Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding
Eun Woo Im, Dhruv Madhwal, Vivek Gupta
Vision-Language Models demonstrate remarkable capabilities but often struggle with compositional reasoning, exhibiting vulnerabilities regarding word order and attribute binding. T…
cs.CV2026
Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
Eun Woo Im, Muhammad Kashif Ali, Vivek Gupta
Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal capabilities, but they inherit the tendency to hallucinate from their underlying language models. While…
cs.CV2025
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
Chaoyu Li, Eun Woo Im, Pooyan Fazli
Multimodal large language models (MLLMs) have recently shown significant advancements in video understanding, excelling in content reasoning and instruction-following tasks. Howeve…