3 papers
cs.CV2026
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
cs.LG2026
Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions
Nicole Ma, Nick Rui
We study planning site formation in language models -- where internal representations of structurally-constrained future tokens form during the forward pass, and whether they causa…
cs.RO2026
Openpi Comet: Competition Solution For 2025 BEHAVIOR Challenge
Junjie Bai, Yu-Wei Chao, Qizhi Chen +13
The 2025 BEHAVIOR Challenge is designed to rigorously track progress toward solving long-horizon tasks by physical agents in simulated environments. BEHAVIOR-1K focuses on everyday…