6 papers · 1 filter
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers
Trung X. Pham, Kang Zhang, Ji Woo Hong +1
Diffusion Transformers have achieved state-of-the-art performance in class-conditional and multimodal generation, yet the structure of their learned conditional embeddings remains…
World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2…
ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On
Ji Woo Hong, Tri Ton, Trung X. Pham +3
This paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of pr…
E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization
Trung X. Pham, Zhang Kang, Ji Woo Hong +2
We propose E-MD3C (fficient asked iffusion Transformer with Disentangled onditions and ompact $\underline…
Cross-view Masked Diffusion Transformers for Person Image Synthesis
Trung X. Pham, Zhang Kang, Chang D. Yoo
We present X-MDPT (-view asked iffusion rediction ransformers), a novel diffusion model designed for…