activity
20232025
most citedMALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20251 cited

MALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation

Sihyun Yu, Meera Hahn, Dan Kondratyuk +6

Diffusion models are successful for synthesizing high-quality videos but are limited to generating short clips (e.g., 2-10 seconds). Synthesizing sustained footage (e.g. over minut…

cs.CV2024

Learning Complex Non-Rigid Image Edits from Multimodal Conditioning

Nikolai Warner, Jack Kolb, Meera Hahn +3

In this paper we focus on inserting a given human (specifically, a single image of a person) into a novel scene. Our method, which builds on top of Stable Diffusion, yields natural…

cs.CV2023

Photorealistic Video Generation with Diffusion Models

Agrim Gupta, Lijun Yu, Kihyuk Sohn +6

We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach has two key design decisions. First, we use a causal encod…

cs.CV2023

VideoPoet: A Large Language Model for Zero-Shot Video Generation

Dan Kondratyuk, Lijun Yu, Xiuye Gu +28

We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-on…

cs.CV2023

Text and Click inputs for unambiguous open vocabulary instance segmentation

Nikolai Warner, Meera Hahn, Jonathan Huang +2

Segmentation localizes objects in an image on a fine-grained per-pixel scale. Segmentation benefits by humans-in-the-loop to provide additional input of objects to segment using a…