activity
20192021
most citedAre we pretraining it right? Digging deeper into visio-linguistic pretraining

30 citations · 73 across the 4 of their papers we have counts for

collaborators

8 papers

cs.CL2021

Tricks for Training Sparse Translation Models

Dheeru Dua, Shruti Bhosale, Vedanuj Goswami +3

Multi-task learning with an unbalanced data distribution skews model learning towards high resource tasks, especially when model capacity is fixed and fully shared across all tasks…

cs.CV202122 cited

Human-Adversarial Visual Question Answering

Sasha Sheng, Amanpreet Singh, Vedanuj Goswami +4

Performance on the most commonly used Visual Question Answering dataset (VQA v2) is starting to approach human accuracy. However, in interacting with state-of-the-art VQA models, i…

cs.CV202021 cited

Creative Sketch Generation

Songwei Ge, Vedanuj Goswami, C. Lawrence Zitnick +1

Sketching or doodling is a popular creative activity that people engage in. However, most existing work in automatic sketch understanding or generation has focused on sketches that…

cs.AI2020

The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes

Douwe Kiela, Hamed Firooz, Aravind Mohan +4

This work proposes a new challenge set for multimodal classification, focusing on detecting hate speech in multimodal memes. It is constructed such that unimodal models struggle an…

cs.CV202030 cited

Are we pretraining it right? Digging deeper into visio-linguistic pretraining

Amanpreet Singh, Vedanuj Goswami, Devi Parikh

Numerous recent works have proposed pretraining generic visio-linguistic representations and then finetuning them for downstream vision and language tasks. While architecture and o…

cs.CV2020

MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond

Duy-Kien Nguyen, Vedanuj Goswami, Xinlei Chen

This paper focuses on visual counting, which aims to predict the number of occurrences given a natural image and a query (e.g. a question or a category). Unlike most prior works th…