3 papers
cs.LG2026
Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding
Eun Woo Im, Dhruv Madhwal, Vivek Gupta
Vision-Language Models demonstrate remarkable capabilities but often struggle with compositional reasoning, exhibiting vulnerabilities regarding word order and attribute binding. T…
cs.CV2026
Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
Eun Woo Im, Muhammad Kashif Ali, Vivek Gupta
Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal capabilities, but they inherit the tendency to hallucinate from their underlying language models. While…
cs.CV2025
Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization
Muhammad Kashif Ali, Eun Woo Im, Dongjin Kim +4
Video stabilization remains a fundamental problem in computer vision, particularly pixel-level synthesis solutions for video stabilization, which synthesize full-frame outputs, add…