3 papers
cs.CV2026
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
cs.CV2026
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
Sethuraman T, Savya Khosla, Aditi Tiwari +11
This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, su…
cs.CV2025
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
Sethuraman TV, Savya Khosla, Vignesh Srinivasakumar +5
Dense video prediction tasks, such as object tracking and semantic segmentation, require video encoders that generate temporally consistent, spatially dense features for every fram…