29 citations · 76 across the 9 of their papers we have counts for
4 papers · 1 filter
An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding
Wei Chen, Long Chen, Yu Wu
Most advanced visual grounding methods rely on Transformers for visual-linguistic feature fusion. However, these Transformer-based approaches encounter a significant drawback: the…
RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments
Mengxue Qu, Yu Wu, Wu Liu +4
Intention-oriented object detection aims to detect desired objects based on specific intentions or requirements. For instance, when we desire to "lie down and rest", we instinctive…
Click-Feedback Retrieval
Zeyu Wang, Yu Wu
Retrieving target information based on input query is of fundamental importance in many real-world applications. In practice, it is not uncommon for the initial search to fail, whe…
SiRi: A Simple Selective Retraining Mechanism for Transformer-based Visual Grounding
Mengxue Qu, Yu Wu, Wu Liu +5
In this paper, we investigate how to achieve better visual grounding with modern vision-language transformers, and propose a simple yet powerful Selective Retraining (SiRi) mechani…