activity
20242026
collaborators

6 papers

cs.CV2026

MobileWan: Closing the Quality Gap for Mobile Video Diffusion

Mohsen Ghafoorian, Denis Korzhenkov, Adil Karjauv +9

Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual fidelity and motion coheren…

cs.CV2026

Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

Long Vu, Tan Ngo, Animesh Karnewar +5

Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and…

cs.CV2026

PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference

Denis Korzhenkov, Adil Karjauv, Animesh Karnewar +2

Recently proposed pyramidal models decompose the conventional forward and backward diffusion processes into multiple stages operating at varying resolutions. These models handle in…

cs.CV2025

Neodragon: Mobile Video Generation using Diffusion Transformer

Animesh Karnewar, Denis Korzhenkov, Ioannis Lelekas +10

We introduce Neodragon, a text-to-video system capable of generating 2s (49 frames @24 fps) videos at the 640x1024 resolution directly on a Qualcomm Hexagon NPU in a record 6.7s (7…

cs.CV2024

GOEmbed: Gradient Origin Embeddings for Representation Agnostic 3D Feature Learning

Animesh Karnewar, Roman Shapovalov, Tom Monnier +3

Encoding information from 2D views of an object into a 3D representation is crucial for generalized 3D feature extraction. Such features can then enable 3D reconstruction, 3D gener…

cs.CV2024

Meta 3D Gen

Raphael Bensadoun, Tom Monnier, Yanir Kleiman +17

We introduce Meta 3D Gen (3DGen), a new state-of-the-art, fast pipeline for text-to-3D asset generation. 3DGen offers 3D asset creation with high prompt fidelity and high-quality 3…