4 citations · 4 across the 10 of their papers we have counts for
11 papers
Spot-the-shift: Evaluating Grounded Image Difference Captioning of Long-term Changes
Benedetta Liberatori, Nermin Samet, Paolo Rota +4
Long-term change understanding from images of the same place revisited over time is a challenging task with applications in map maintenance and urban infrastructure monitoring. Pri…
Boosting Visual Instruction Tuning with Self-Supervised Guidance
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc +2
Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-grained visual reasoning. Rece…
Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models
Katarzyna Zaleska, Łukasz Popek, Monika Wysoczańska +1
Text-to-image diffusion models exhibit remarkable generative capabilities, yet their internal operations remain opaque, particularly when handling prompts that are not fully descri…
Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models
Monika Wysoczańska, Antonin Vobecky, Amaia Cardiel +4
Recent CLIP-like Vision-Language Models (VLMs), pre-trained on large amounts of image-text pairs to align both modalities with a simple contrastive objective, have paved the way to…
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
Monika Wysoczańska, Oriane Siméoni, Michaël Ramamonjisoa +3
The popular CLIP model displays impressive zero-shot capabilities thanks to its seamless interaction with arbitrary text prompts. However, its lack of spatial awareness makes it un…
Tell Me What Is Good About This Property: Leveraging Reviews For Segment-Personalized Image Collection Summarization
Monika Wysoczanska, Moran Beladev, Karen Lastmann Assaraf +4
Image collection summarization techniques aim to present a compact representation of an image gallery through a carefully selected subset of images that captures its semantic conte…