4 papers
Detect Anything via Next Point Prediction
Qing Jiang, Junan Huo, Xingyu Chen +6
Object detection has long been dominated by traditional coordinate regression-based models, such as YOLO, DETR, and Grounding DINO. Although recent efforts have attempted to levera…
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
Tianhe Ren, Yihao Chen, Qing Jiang +17
In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X…
Referring to Any Person
Qing Jiang, Lin Wu, Zhaoyang Zeng +5
Humans are undoubtedly the most important participants in computer vision, and the ability to detect any individual given a natural language description, a task we define as referr…
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Qing Jiang, Gen Luo, Yuqin Yang +5
Perception and understanding are two pillars of computer vision. While multimodal large language models (MLLM) have demonstrated remarkable visual understanding capabilities, they…