3 papers
cs.CV2026
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
Tripti Shukla, Zsolt Kira
Large vision-language models (VLMs) frequently suffer from hallucinations, generating content that is inconsistent with visual inputs. Existing methods typically address this probl…
cs.CL2025
Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback
Ananth Muppidi, Tarak Das, Sambaran Bandyopadhyay +2
The generation of presentation slides automatically is an important problem in the era of generative AI. This paper focuses on evaluating multimodal content in presentation slides…
cs.CV2024
Design-o-meter: Towards Evaluating and Refining Graphic Designs
Sahil Goyal, Abhinav Mahajan, Swasti Mishra +4
Graphic designs are an effective medium for visual communication. They range from greeting cards to corporate flyers and beyond. Off-late, machine learning techniques are able to g…