3 papers
cs.CV2026
Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding
Xiuyuan Zhu, Ke Lu, Kun Dong +6
Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify thi…
cs.CV2026
DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models
Xin Jiang, Hao Tang, Junyao Gao +4
Recent advances in generative models have shifted AI-generated image detection from identifying easily distinguishable, fully synthetic images to identifying highly realistic conte…
cs.CV2026
IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models
Xiuyuan Zhu, Ke Lu, Hao Wu +4
Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs bounding-box coordinates as text given…