Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
Tripti Shukla, Zsolt Kira
Large vision-language models (VLMs) frequently suffer from hallucinations, generating content that is inconsistent with visual inputs. Existing methods typically address this probl…
cs.CV2024
Design-o-meter: Towards Evaluating and Refining Graphic Designs
Sahil Goyal, Abhinav Mahajan, Swasti Mishra +4
Graphic designs are an effective medium for visual communication. They range from greeting cards to corporate flyers and beyond. Off-late, machine learning techniques are able to g…
cs.CV2024
Test-time Conditional Text-to-Image Synthesis Using Diffusion Models
Tripti Shukla, Srikrishna Karanam, Balaji Vasan Srinivasan
We consider the problem of conditional text-to-image synthesis with diffusion models. Most recent works need to either finetune specific parts of the base diffusion model or introd…