6 citations · 10 across the 19 of their papers we have counts for
32 papers · 1 filter
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Xin Zhou, Zongchuang Zhao, Zhibo Yang +13
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-langu…
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
Zongchuang Zhao, Xin Zhou, Tianyang Xu +6
World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods incur costly test-time future imag…
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Hengyi Xie, Chenfei Yao, Xianjin Wu +4
Vision-language-action (VLA) models commonly adopt an LLM-centric pathway, where visual observations are projected into the representation space of a large language…
Wat3R: Underwater 3D Geometry Learning without Annotations
Jiangwei Ren, Xingyu Jiang, Zijie Song +4
Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pion…
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
Xin Zhou, Dingkang Liang, Xiwu Chen +4
Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominantly focus on future scene gen…
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
Zhengyang Sun, Yu Chen, Xin Zhou +4
Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. We introduce NUMINA…