11 papers
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
Vesta: A Generalist Embodied Reasoning Model
Johan Bjorck, Zhiqi Li, Yunze Man +29
Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at indiv…
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
Shihao Wang, Guo Chen, De-an Huang +6
While Video Large Language Models (Video-LLMs) have shown significant potential in multimodal understanding and reasoning tasks, how to efficiently select the most informative fram…
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
Guo Chen, Lidong Lu, Yicheng Liu +17
While datasets for video understanding have scaled to hour-long durations, they typically consist of densely concatenated clips that differ from natural, unscripted daily life. To…
LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
Yunze Man, Shihao Wang, Guowen Zhang +7
To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models (VLMs) excel at open-ended 2D description and grounding, yet multi-ob…
Language Models are Symbolic Learners in Arithmetic
Chunyuan Deng, Zhiqi Li, Roy Xie +2
The prevailing question in LM performing arithmetic is whether these models learn to truly compute or if they simply master superficial pattern matching. In this paper, we argues f…