Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
cs.CV2026
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
Baifeng Shi, Stephanie Fu, Long Lian +10
Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in th…
cs.CV2024
Synthesizing Moving People with 3D Control
Boyi Li, Junming Chen, Jathushan Rajasegaran +3
In this paper, we present a diffusion model-based framework for animating people from a single image for a given target 3D motion sequence. Our approach has two core components: a)…