4 papers
PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
Denis Korzhenkov, Adil Karjauv, Animesh Karnewar +2
Recently proposed pyramidal models decompose the conventional forward and backward diffusion processes into multiple stages operating at varying resolutions. These models handle in…
Neodragon: Mobile Video Generation using Diffusion Transformer
Animesh Karnewar, Denis Korzhenkov, Ioannis Lelekas +10
We introduce Neodragon, a text-to-video system capable of generating 2s (49 frames @24 fps) videos at the 640x1024 resolution directly on a Qualcomm Hexagon NPU in a record 6.7s (7…
3inGAN: Learning a 3D Generative Model from Images of a Self-similar Scene
Animesh Karnewar, Oliver Wang, Tobias Ritschel +1
We introduce 3inGAN, an unconditional 3D generative model trained from 2D images of a single self-similar 3D scene. Such a model can be used to produce 3D "remixes" of a given scen…
MSG-GAN: Multi-Scale Gradients for Generative Adversarial Networks
Animesh Karnewar, Oliver Wang
While Generative Adversarial Networks (GANs) have seen huge successes in image synthesis tasks, they are notoriously difficult to adapt to different datasets, in part due to instab…