7 citations · 22 across the 12 of their papers we have counts for
3 papers · 2 filters
A Multi-level Alignment Training Scheme for Video-and-Language Grounding
Yubo Zhang, Feiyang Niu, Qing Ping +1
To solve video-and-language grounding tasks, the key is for the network to understand the connection between the two modalities. For a pair of video and language description, their…
Privacy Preserving Visual Question Answering
Cristian-Paul Bara, Qing Ping, Abhinav Mathur +3
We introduce a novel privacy-preserving methodology for performing Visual Question Answering on the edge. Our method constructs a symbolic representation of the visual scene, using…
A Thousand Words Are Worth More Than a Picture: Natural Language-Centric Outside-Knowledge Visual Question Answering
Feng Gao, Qing Ping, Govind Thattai +3
Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information…