From the 1 of 7 linked papers with an AI index.
4 papers · 1 filter
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
Parsa Hosseini, Sumit Nawathe, Mazda Moayeri +2
The paper introduces SpurLens, an automated pipeline that uses GPT-4 and open-set object detectors to find spurious visual cues in multimodal large language models, showing that th…
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
Arman Zarei, Keivan Rezaei, Samyadeep Basu +4
Text-to-image diffusion-based generative models have the stunning ability to generate photo-realistic images and achieve state-of-the-art low FID scores on challenging image genera…
Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class
Mazda Moayeri, Michael Rabbat, Mark Ibrahim +1
Vision-language models enable open-world classification of objects without the need for any retraining. While this zero-shot paradigm marks a significant advance, even today's best…
Rethinking Artistic Copyright Infringements in the Era of Text-to-Image Generative Models
Mazda Moayeri, Samyadeep Basu, Sriram Balasubramanian +4
Recent text-to-image generative models such as Stable Diffusion are extremely adept at mimicking and generating copyrighted content, raising concerns amongst artists that their uni…