6 papers
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
Debottam Dutta, Jaehoon Hahm, Jianchong Chen +1
Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images m…
Steer Away From Mode Collisions: Improving Composition In Diffusion Models
Debottam Dutta, Jianchong Chen, Rajalaxmi Rajagopalan +2
We propose to improve multi-concept prompt fidelity in text-to-image diffusion models. We begin with common failure cases - prompts like "a cat and a dog" that sometimes yields ima…
Personalized Image Generation via Human-in-the-loop Bayesian Optimization
Rajalaxmi Rajagopalan, Debottam Dutta, Yu-Lin Wei +1
Imagine Alice has a specific image in her mind, say, the view of the street in which she grew up during her childhood. To generate that exact image, she guides a generativ…
Learning Energy-based Variational Latent Prior for VAEs
Debottam Dutta, Chaitanya Amballa, Zhongweiyang Xu +2
Variational Auto-Encoders (VAEs) are known to generate blurry and inconsistent samples. One reason for this is the "prior hole" problem. A prior hole refers to regions that have hi…
Multi-Source Music Generation with Latent Diffusion
Zhongweiyang Xu, Debottam Dutta, Yu-Lin Wei +1
Most music generation models directly generate a single music mixture. To allow for more flexible and controllable generation, the Multi-Source Diffusion Model (MSDM) has been prop…
Estimating Multi-chirp Parameters using Curvature-guided Langevin Monte Carlo
Sattwik Basu, Debottam Dutta, Yu-Lin Wei +1
This paper considers the problem of estimating chirp parameters from a noisy mixture of chirps. While a rich body of work exists in this area, challenges remain when extending thes…