7 papers
Localizing Knowledge in Diffusion Transformers
Arman Zarei, Samyadeep Basu, Keivan Rezaei +3
Understanding how knowledge is distributed across the layers of generative models is crucial for improving interpretability, controllability, and adaptation. While prior work has e…
AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
Arman Zarei, Jiacheng Pan, Matthew Gwilliam +2
Text-to-image generative models have achieved remarkable visual quality but still struggle with compositionalityaccurately capturing object relationships, attribute bindings, an…
SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
Arman Zarei, Samyadeep Basu, Mobina Pournemat +3
Instruction-based image editing models have recently achieved impressive performance, enabling complex edits to an input image from a multi-instruction prompt. However, these model…
Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
Mobina Pournemat, Keivan Rezaei, Gaurang Sriramanan +5
Despite widespread success in language understanding and generation, large language models (LLMs) exhibit unclear and often inconsistent behavior when faced with tasks that require…
Understanding the Effect of using Semantically Meaningful Tokens for Visual Representation Learning
Neha Kalibhat, Priyatham Kattakinda, Sumit Nawathe +5
Vision transformers have established a precedent of patchifying images into uniformly-sized chunks before processing. We hypothesize that this design choice may limit models in lea…
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
Arman Zarei, Keivan Rezaei, Samyadeep Basu +4
Text-to-image diffusion-based generative models have the stunning ability to generate photo-realistic images and achieve state-of-the-art low FID scores on challenging image genera…