1 citations · 3 across the 13 of their papers we have counts for
Showing 2026 · cs.CVShow all
2 papers · 2 filters
cs.CV2026
Token-Based Affordance Grounding with Large Vision-Language Models
Seung Il Lee, Qinqian Lei, Daguang Xu +4
Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence and embodied perception. Previous studies…
cs.CV2026
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…