5 citations · 5 across the 2 of their papers we have counts for
3 papers
cs.CV2022★ 5 cited
A Thousand Words Are Worth More Than a Picture: Natural Language-Centric Outside-Knowledge Visual Question Answering
Feng Gao, Qing Ping, Govind Thattai +3
Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information…
cs.CV2021
BioFors: A Large Biomedical Image Forensics Dataset
Ekraam Sabir, Soumyaroop Nandi, Wael AbdAlmageed +1
Research in media forensics has gained traction to combat the spread of misinformation. However, most of this research has been directed towards content generated on social media.…
cs.CV2021
Learning Better Visual Dialog Agents with Pretrained Visual-Linguistic Representation
Tao Tu, Qing Ping, Govind Thattai +2
GuessWhat?! is a two-player visual dialog guessing game where player A asks a sequence of yes/no questions (Questioner) and makes a final guess (Guesser) about a target object in a…