Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Reflective VLA: In-Context Action Consequences Make VLAs Generalize
Qing Lian, Kent Yu, Lei Zhang
Most vision-language-action (VLA) models are reactive: they predict the next action from the current instruction and observation, implicitly assuming that the current observation f…
cs.CV2025
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
Tianhe Ren, Yihao Chen, Qing Jiang +17
In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X…
cs.CV2024
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Tianhe Ren, Qing Jiang, Shilong Liu +13
This paper introduces Grounding DINO 1.5, a suite of advanced open-set object detection models developed by IDEA Research, which aims to advance the "Edge" of open-set object detec…