5 papers
VPTracker: Global Vision-Language Tracking via Visual Prompt
Jingchao Wang, Kaiwen Zhou, Zhijian Wu +3
Vision-Language Tracking aims to continuously localize objects described by a visual template and a language description. Existing methods, however, are typically limited to local…
A Simple and Better Baseline for Visual Grounding
Jingchao Wang, Wenlong Zhang, Dingjiang Huang +2
Visual grounding aims to predict the locations of target objects specified by textual descriptions. For this task with linguistic and visual modalities, there is a latest research…
Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
Jingchao Wang, Zhijian Wu, Dingjiang Huang +2
Reference Expression Segmentation (RES) aims to segment image regions specified by referring expressions and has become popular with the rise of multimodal large models (MLLMs). Wh…
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
Jingchao Wang, Hong Wang, Wenlong Zhang +3
Multi-task visual grounding (MTVG) includes two sub-tasks, i.e., Referring Expression Comprehension (REC) and Referring Expression Segmentation (RES). The existing representative a…
Imagination-Limited Q-Learning for Offline Reinforcement Learning
Wenhui Liu, Zhijian Wu, Jingchao Wang +2
Offline reinforcement learning seeks to derive improved policies entirely from historical data but often struggles with over-optimistic value estimates for out-of-distribution (OOD…