4 papers
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2…
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
Andrew Z. Wang, Songwei Ge, Tero Karras +2
Both text-to-image generation and large language models (LLMs) have made significant advancements. However, many text-to-image models still employ the somewhat outdated T5 and CLIP…
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
NVIDIA, :, Alisson Azzolini +51
Physical AI systems need to perceive, understand, and perform complex actions in the physical world. In this paper, we present the Cosmos-Reason1 models that can understand the phy…