activity
20242026
most citedCosmos World Foundation Model Platform for Physical AI

11 citations · 16 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2026

Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation

Yifan Li, Xinyu Zhou, Yunhao Ge +2

Language instructions offer a flexible way to specify task semantics but can become ambiguous in cluttered scenes containing multiple visually similar objects and candidate targets…

cs.CV2026

VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning

Chaoyang Wang, Wenrui Bao, Sicheng Gao +5

Vision-Language-Action (VLA) models have shown promising capabilities for embodied intelligence, but most existing approaches rely on text-based chain-of-thought reasoning where vi…

cs.CV2026

I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners

Lu Ling, Yunhao Ge, Yichen Sheng +1

Generalization remains the central challenge for interactive 3D scene generation. Existing learning-based approaches ground spatial understanding in limited scene dataset, restrict…

cs.CV2025

World Simulation with Video Foundation Models for Physical AI

NVIDIA, :, Arslan Ali +87

We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2…

cs.CV2025

Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding

Haotian Xue, Yunhao Ge, Yu Zeng +4

Vision-Language Models (VLMs) have demonstrated impressive world knowledge across a wide range of tasks, making them promising candidates for embodied reasoning applications. Howev…

cs.CV2025

ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary

Zeqi Gu, Yin Cui, Zhaoshuo Li +6

Designing 3D scenes is traditionally a challenging task that demands both artistic expertise and proficiency with complex software. Recent advances in text-to-3D generation have gr…