activity
20242026
collaborators

15 papers

cs.RO2026

Vesta: A Generalist Embodied Reasoning Model

Johan Bjorck, Zhiqi Li, Yunze Man +29

Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at indiv…

cs.RO2026

Planning-aligned Token Compression for Long-Context Autonomous Driving

Zhixuan Liang, Yuxiao Chen, Yurong You +12

Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences that quickly exceed real-time computationa…

cs.LG2026

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

NVIDIA, :, Amala Sanjay Deshmukh +204

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 N…

cs.RO2026

Accelerating Structured Chain-of-Thought in Autonomous Vehicles

Yi Gu, Yan Wang, Yuxiao Chen +8

Chain-of-Thought (CoT) reasoning enhances the decision-making capabilities of vision-language-action models in autonomous driving, but its autoregressive nature introduces signific…

cs.RO2025

Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning

Zhenghao "Mark" Peng, Wenhao Ding, Yurong You +11

Recent reasoning-augmented Vision-Language-Action (VLA) models have improved the interpretability of end-to-end autonomous driving by generating intermediate reasoning traces. Yet…

cs.CV2025

Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving

Jiawei Yang, Ziyu Chen, Yurong You +7

We present Flex, an efficient and effective scene encoder that addresses the computational bottleneck of processing high-volume multi-camera data in end-to-end autonomous driving.…