6 papers
Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation
Matthew Bendel, Stephen W. Bailey, Mithilesh Vaidya +2
Long-horizon video generation suffers from two intertwined issues. First, there is drift, where video quality degrades over time. Second, there are continuity issues which manifest…
PoDAR: Power-Disentangled Audio Representation for Generative Modeling
Alejandro Luebs, Mithilesh Vaidya, Ishaan Kumar +5
The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focu…
DreamTexture: Shape from Virtual Texture with Analysis by Augmentation
Ananta R. Bhattarai, Xingzhe He, Alla Sheffer +1
DreamFusion established a new paradigm for unsupervised 3D reconstruction from virtual views by combining advances in generative models and differentiable rendering. However, the u…
A Data Perspective on Enhanced Identity Preservation for Diffusion Personalization
Xingzhe He, Zhiwen Cao, Nicholas Kolkin +4
Large text-to-image models have revolutionized the ability to generate imagery using natural language. However, particularly unique or personal visual concepts, such as pets and fu…
LatentKeypointGAN: Controlling Images via Latent Keypoints
Xingzhe He, Bastian Wandt, Helge Rhodin
Generative adversarial networks (GANs) have attained photo-realistic quality in image generation. However, how to best control the image content remains an open challenge. We intro…
Unsupervised Keypoints from Pretrained Diffusion Models
Eric Hedlin, Gopal Sharma, Shweta Mahajan +5
Unsupervised learning of keypoints and landmarks has seen significant progress with the help of modern neural network architectures, but performance is yet to match the supervised…