6 papers
Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
Mayank Vatsa, Aparna Bharati, Richa Singh
The architectural blueprint of today's leading text-to-image models contains a fundamental flaw: an inability to handle logical composition. This survey investigates this breakdown…
StyleProtect: Safeguarding Artistic Identity in Fine-tuned Diffusion Models
Qiuyu Tang, Joshua Krinsky, Aparna Bharati
The rapid advancement of generative models, particularly diffusion-based approaches, has inadvertently facilitated their potential for misuse. Such models enable malicious exploite…
Is Perturbation-Based Image Protection Disruptive to Image Editing?
Qiuyu Tang, Bonor Ayambem, Mooi Choo Chuah +1
The remarkable image generation capabilities of state-of-the-art diffusion models, such as Stable Diffusion, can also be misused to spread misinformation and plagiarize copyrighted…
Exploring Saliency Bias in Manipulation Detection
Joshua Krinsky, Alan Bettis, Qiuyu Tang +2
The social media-fuelled explosion of fake news and misinformation supported by tampered images has led to growth in the development of models and datasets for image manipulation d…
Subjective Face Transform using Human First Impressions
Chaitanya Roygaga, Joshua Krinsky, Kai Zhang +2
Humans tend to form quick subjective first impressions of non-physical attributes when seeing someone's face, such as perceived trustworthiness or attractiveness. To understand wha…
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
Mayank Vatsa, Aparna Bharati, Surbhi Mittal +1
Negation, a linguistic construct conveying absence, denial, or contradiction, poses significant challenges for multilingual multimodal foundation models. These models excel in task…