16 papers · 1 filter
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
Tao Lin, Yuxin Du, Jiting Liu +14
Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often s…
CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers
Zhuojin Li, Hsin-Pai Cheng, Hong Cai +2
Diffusion Transformers (DiTs) deliver remarkable image and video generation quality but incur high computational cost, limiting scalability and on-device deployment. We introduce C…
Generative Scenario Rollouts for End-to-End Autonomous Driving
Rajeev Yasarla, Deepti Hegde, Shizhong Han +10
Vision-Language-Action (VLA) models are emerging as highly effective planning models for end-to-end autonomous driving systems. However, current works mostly rely on imitation lear…
ODG: Occupancy Prediction Using Dual Gaussians
Yunxiao Shi, Yinhao Zhu, Shizhong Han +4
Occupancy prediction infers fine-grained 3D geometry and semantics from camera images of the surrounding environment, making it a critical perception task for autonomous driving. E…
DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos
Rajeev Yasarla, Shizhong Han, Hong Cai +1
Camera-based 3D object detection in Bird's Eye View (BEV) is one of the most important perception tasks in autonomous driving. Earlier methods rely on dense BEV features, which are…
RoCA: Robust Cross-Domain End-to-End Autonomous Driving
Rajeev Yasarla, Shizhong Han, Hsin-Pai Cheng +7
End-to-end (E2E) autonomous driving has recently emerged as a new paradigm, offering significant potential. However, few studies have looked into the practical challenge of deploym…