Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
Michael Ogezi, Freda Shi
Vision-language models (VLMs) work well in tasks ranging from image captioning to visual question answering (VQA), yet they struggle with spatial reasoning, a key skill for underst…
cs.CV2024
Semantically-Prompted Language Models Improve Visual Descriptions
Michael Ogezi, Bradley Hauer, Grzegorz Kondrak
Language-vision models like CLIP have made significant strides in vision tasks, such as zero-shot image classification (ZSIC). However, generating specific and expressive visual de…
cs.CV2024
Optimizing Negative Prompts for Enhanced Aesthetics and Fidelity in Text-To-Image Generation
Michael Ogezi, Ning Shi
In text-to-image generation, using negative prompts, which describe undesirable image characteristics, can significantly boost image quality. However, producing good negative promp…