11 citations · 16 across the 16 of their papers we have counts for
14 papers · 1 filter
Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation
Yifan Li, Xinyu Zhou, Yunhao Ge +2
Language instructions offer a flexible way to specify task semantics but can become ambiguous in cluttered scenes containing multiple visually similar objects and candidate targets…
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
Chaoyang Wang, Wenrui Bao, Sicheng Gao +5
Vision-Language-Action (VLA) models have shown promising capabilities for embodied intelligence, but most existing approaches rely on text-based chain-of-thought reasoning where vi…
I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners
Lu Ling, Yunhao Ge, Yichen Sheng +1
Generalization remains the central challenge for interactive 3D scene generation. Existing learning-based approaches ground spatial understanding in limited scene dataset, restrict…
World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2…
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
Haotian Xue, Yunhao Ge, Yu Zeng +4
Vision-Language Models (VLMs) have demonstrated impressive world knowledge across a wide range of tasks, making them promising candidates for embodied reasoning applications. Howev…
ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
Zeqi Gu, Yin Cui, Zhaoshuo Li +6
Designing 3D scenes is traditionally a challenging task that demands both artistic expertise and proficiency with complex software. Recent advances in text-to-3D generation have gr…