10 papers
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling
Tao Xu, Runhao Zhang, Zhijian Huang +7
Occluded tasks remain a bottleneck in robot manipulation. Existing solutions either deploy additional physical cameras requiring training-inference camera parity, or rely on explic…
Deformba: Vision State Space Model with Adaptive State Fusion
Hongyu Ke, Jack Morris, Yongkang Liu +4
State Space Models (SSMs) have emerged as a powerful and efficient alternative to Transformers, demonstrating linear-time complexity and exceptional sequence modeling capabilities.…
DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models
Ye Sun, Xin Wang, Jiaming Zhang +7
While vision and multimodal foundation models underpin critical tasks from perception to complex reasoning, they remain highly vulnerable to adversarial attacks. However, tradition…
InEdit-Bench: Benchmarking Intermediate Logical Pathways for Intelligent Image Editing Models
Zhiqiang Sheng, Xumeng Han, Zhiwei Zhang +6
Multimodal generative models have made significant strides in image editing, demonstrating impressive performance on a variety of static tasks. However, their proficiency typically…
World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2…