6 papers
Steer Away From Mode Collisions: Improving Composition In Diffusion Models
Debottam Dutta, Jianchong Chen, Rajalaxmi Rajagopalan +2
We propose to improve multi-concept prompt fidelity in text-to-image diffusion models. We begin with common failure cases - prompts like "a cat and a dog" that sometimes yields ima…
Can NeRFs See without Cameras?
Chaitanya Amballa, Sattwik Basu, Yu-Lin Wei +3
Neural Radiance Fields (NeRFs) have been remarkably successful at synthesizing novel views of 3D scenes by optimizing a volumetric scene function. This scene function models how op…
Learning Energy-based Variational Latent Prior for VAEs
Debottam Dutta, Chaitanya Amballa, Zhongweiyang Xu +2
Variational Auto-Encoders (VAEs) are known to generate blurry and inconsistent samples. One reason for this is the "prior hole" problem. A prior hole refers to regions that have hi…
PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations
Yu Wei, Jiahui Zhang, Xiaoqin Zhang +2
COLMAP-free 3D Gaussian Splatting (3D-GS) has recently attracted increasing attention due to its remarkable performance in reconstructing high-quality 3D scenes from unposed images…
Multi-Source Music Generation with Latent Diffusion
Zhongweiyang Xu, Debottam Dutta, Yu-Lin Wei +1
Most music generation models directly generate a single music mixture. To allow for more flexible and controllable generation, the Multi-Source Diffusion Model (MSDM) has been prop…
Estimating Multi-chirp Parameters using Curvature-guided Langevin Monte Carlo
Sattwik Basu, Debottam Dutta, Yu-Lin Wei +1
This paper considers the problem of estimating chirp parameters from a noisy mixture of chirps. While a rich body of work exists in this area, challenges remain when extending thes…