5 papers
A Unified 3D Object Perception Framework for Real-Time Outside-In Multi-Camera Systems
Yizhou Wang, Sameer Pusegaonkar, Yuxing Wang +10
Accurate 3D object perception and multi-target multi-camera (MTMC) tracking are fundamental for the digital transformation of industrial infrastructure. However, transitioning "ins…
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
Anqi Li, Zhiyong Wang, Jiazhao Zhang +5
Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This ta…
Embodied Navigation Foundation Model
Jiazhao Zhang, Anqi Li, Yunpeng Qi +14
Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments following language instructions.…
Cosmos World Foundation Model Platform for Physical AI
NVIDIA, :, Niket Agarwal +76
Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model. In this paper, we present th…
TrackVLA: Embodied Visual Tracking in the Wild
Shaoan Wang, Jiazhao Zhang, Minghan Li +7
Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision. This task is inh…