4 papers
From Web to Pixels: Bringing Agentic Search into Visual Perception
Bokang Yang, Xinyi Sun, Kaituo Feng +3
Visual perception connects high-level semantic understanding to pixel-level perception, but most existing settings assume that the decisive evidence for identifying a target is alr…
Towards Visual Query Localization in the 3D World
Liang Peng, Bohan Tan, Zhipeng Zhang +4
Visual query localization (VQL) aims to predict the spatio-temporal response of the most recent occurrence in a sequence given a query. Currently, most research focuses on visual q…
DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
Yani Zhang, Dongming Wu, Hao Shi +3
A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pair a detector with a separate gr…
Bootstrapping Referring Multi-Object Tracking
Yani Zhang, Dongming Wu, Wencheng Han +1
Referring understanding is a fundamental task that bridges natural language and visual content by localizing objects described in free-form expressions. However, existing works are…