10 papers
Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
Ryan Ramos, Vladan StojniÄ, Giorgos Kordopatis-Zilos +3
Prior work has analyzed the robustness of visual encoders to image transformations and corruptions, particularly in cases where such alterations are not seen during training. When…
EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
Lu Wei, Yuta Nakashima, Noa Garcia
The widespread adoption of text-to-image (T2I) generation has raised concerns about privacy, bias, and copyright violations. Concept erasure techniques offer a promising solution b…
Privacy in Image Datasets: A Case Study on Pregnancy Ultrasounds
Rawisara Lohanimit, Yankun Wu, Amelia Katirai +2
The rise of generative models has led to increased use of large-scale datasets collected from the internet, often with minimal or no data curation. This raises concerns about the i…
Towards Artwork Explanation in Large-scale Vision Language Models
Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2
Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…
From Global to Local: Social Bias Transfer in CLIP
Ryan Ramos, Yusuke Hirota, Yuta Nakashima +1
The recycling of contrastive language-image pre-trained (CLIP) models as backbones for a large number of downstream tasks calls for a thorough analysis of their transferability imp…
No Annotations for Object Detection in Art through Stable Diffusion
Patrick Ramos, Nicolas Gonthier, Selina Khan +2
Object detection in art is a valuable tool for the digital humanities, as it allows for faster identification of objects in artistic and historical images compared to humans. Howev…