6 citations · 16 across the 5 of their papers we have counts for
7 papers
Improving Selective Visual Question Answering by Learning from Your Peers
Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary +5
Despite advances in Visual Question Answering (VQA), the ability of models to assess their own correctness remains underexplored. Recent work has shown that VQA models, out-of-the-…
EurNet: Efficient Multi-Range Relational Modeling of Spatial Multi-Relational Data
Minghao Xu, Yuanfan Guo, Yi Xu +3
Modeling spatial relationship in the data remains critical across many different tasks, such as image classification, semantic segmentation and protein structure understanding. Pre…
LoopITR: Combining Dual and Cross Encoder Architectures for Image-Text Retrieval
Jie Lei, Xinlei Chen, Ning Zhang +4
Dual encoders and cross encoders have been widely used for image-text retrieval. Between the two, the dual encoder encodes the image and text independently followed by a dot produc…
KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA
Kenneth Marino, Xinlei Chen, Devi Parikh +2
One of the most challenging question types in VQA is when answering the question requires outside knowledge not present in the image. In this work we study open-domain knowledge, t…
Overcoming Statistical Shortcuts for Open-ended Visual Counting
Corentin Dancette, Remi Cadene, Xinlei Chen +1
Machine learning models tend to over-rely on statistical shortcuts. These spurious correlations between parts of the input and the output labels does not hold in real-world setting…
ImVoteNet: Boosting 3D Object Detection in Point Clouds with Image Votes
Charles R. Qi, Xinlei Chen, Or Litany +1
3D object detection has seen quick progress thanks to advances in deep learning on point clouds. A few recent works have even shown state-of-the-art performance with just point clo…