4 papers
Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding
Xiuyuan Zhu, Ke Lu, Kun Dong +6
Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify thi…
IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models
Xiuyuan Zhu, Ke Lu, Hao Wu +4
Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs bounding-box coordinates as text given…
QueryGaussian: Scalable and Training-Free Open-Vocabulary 3D Instance Retrieval
Xiuyuan Zhu, Ke Lu, Zijie Yang +3
Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multimedia analysis. Existing approaches predomi…
Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model
Siwen Jiao, Tianxiong Lv, Kangan Qian +7
Vision-Language Models (VLMs) face a critical bottleneck in achieving precise numerical prediction for 3D scene understanding. Traditional reinforcement learning (RL) approaches, p…