7 papers
Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding
Xiuyuan Zhu, Ke Lu, Kun Dong +6
Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify thi…
IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models
Xiuyuan Zhu, Ke Lu, Hao Wu +4
Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs bounding-box coordinates as text given…
XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
Kangan Qian, ChuChu Xie, Yang Zhong +13
Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud…
Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model
Siwen Jiao, Tianxiong Lv, Kangan Qian +7
Vision-Language Models (VLMs) face a critical bottleneck in achieving precise numerical prediction for 3D scene understanding. Traditional reinforcement learning (RL) approaches, p…
EvaDrive: Evolutionary Adversarial Policy Optimization for End-to-End Autonomous Driving
Siwen Jiao, Kangan Qian, Hao Ye +12
Autonomous driving faces significant challenges in achieving human-like iterative decision-making, which continuously generates, evaluates, and refines trajectory proposals. Curren…
A Survey on Vision-Language-Action Models for Autonomous Driving
Sicong Jiang, Zilin Huang, Kangan Qian +17
The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language unde…