5 papers
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
Yifan Lu, Qi Wu, Jay Zhangjie Wu +4
Most practical high-resolution text-to-image systems, including latent diffusion and autoregressive models, perform generation in a compact latent space, and a decoder maps the gen…
ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
Jay Zhangjie Wu, Xuanchi Ren, Tianchang Shen +11
Recent advances in large generative models have greatly enhanced both image editing and in-context image generation, yet a critical gap remains in ensuring physical consistency, wh…
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
Xuanchi Ren, Yifan Lu, Tianshi Cao +13
Collecting and annotating real-world data for safety-critical physical AI systems, such as Autonomous Vehicle (AV), is time-consuming and costly. It is especially challenging to ca…
Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
NVIDIA, :, Hassan Abu Alhaija +38
We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmen…
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
Xuanchi Ren, Tianchang Shen, Jiahui Huang +7
We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage…