collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

Xiuyuan Zhu, Ke Lu, Kun Dong +6

Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify thi…

cs.CV2026

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models

Xiuyuan Zhu, Ke Lu, Hao Wu +4

Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs bounding-box coordinates as text given…

cs.CV2026

QueryGaussian: Scalable and Training-Free Open-Vocabulary 3D Instance Retrieval

Xiuyuan Zhu, Ke Lu, Zijie Yang +3

Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multimedia analysis. Existing approaches predomi…

cs.CV2024

MambaDETR: Query-based Temporal Modeling using State Space Model for Multi-View 3D Object Detection

Tong Ning, Ke Lu, Xirui Jiang +1

Utilizing temporal information to improve the performance of 3D detection has made great progress recently in the field of autonomous driving. Traditional transformer-based tempora…

cs.CV2024

MVLLaVA: An Intelligent Agent for Unified and Flexible Novel View Synthesis

Hanyu Jiang, Jian Xue, Xing Lan +2

This paper introduces MVLLaVA, an intelligent agent designed for novel view synthesis tasks. MVLLaVA integrates multiple multi-view diffusion models with a large multimodal model,…