3 papers
cs.CV2026
Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models
Katarzyna Zaleska, Åukasz Popek, Monika WysoczaÅska +1
Text-to-image diffusion models exhibit remarkable generative capabilities, yet their internal operations remain opaque, particularly when handling prompts that are not fully descri…
cs.CV2025
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
Monika WysoczaÅska, Shyamal Buch, Anurag Arnab +1
Large vision-language models (VLMs) often struggle to generate long and factual captions. However, traditional measures for hallucination and factuality are not well suited for eva…
cs.CV2025
Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models
Monika WysoczaÅska, Antonin Vobecky, Amaia Cardiel +4
Recent CLIP-like Vision-Language Models (VLMs), pre-trained on large amounts of image-text pairs to align both modalities with a simple contrastive objective, have paved the way to…