6 papers
SLAMFormer-: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing
Zhijian Fang, Weicheng Zheng, Yijun Yuan +7
We introduce the Infinite SLAM Transformer (SLAMFormer-), the first geometric transformer capable of supporting both long-range frontend and backend processing without an e…
DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions
Weicheng Zheng, Yixin Huang, Qiao Sun +2
Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise. We propose DriveMA, a Driv…
Is Your Trajectory Displacement Safe in Long-tail?
Qiao Sun, Weicheng Zheng, Yixin Huang +1
Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing evaluation pipelines are rarely human-aligne…
DriveMA: Rethinking Language Interfaces in Driving VLAs with One-Step Meta-Actions
Weicheng Zheng, Yixin Huang, Qiao Sun +2
Driving Vision-Language-Action Models (Driving VLAs) commonly introduce natural-language reasoning as an intermediate interface for end-to-end planning, but reasoning-centric inter…
DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking
Weicheng Zheng, Xiaofei Mao, Nanfei Ye +4
The advent of Vision-Language Models (VLMs) has significantly advanced end-to-end autonomous driving, demonstrating powerful reasoning abilities for high-level behavior planning ta…
Multi-Representation Adapter with Neural Architecture Search for Efficient Range-Doppler Radar Object Detection
Zhiwei Lin, Weicheng Zheng, Yongtao Wang
Detecting objects efficiently from radar sensors has recently become a popular trend due to their robustness against adverse lighting and weather conditions compared with cameras.…