6 papers
Improving Object Detection and Attribute Recognition by Feature Entanglement Reduction
Zhaoheng Zheng, Arka Sadhu, Ram Nevatia
We explore object detection with two attributes: color and material. The task aims to simultaneously detect objects and infer their color and material. A straight-forward approach…
Video Question Answering with Phrases via Semantic Roles
Arka Sadhu, Kan Chen, Ram Nevatia
Video Question Answering (VidQA) evaluation metrics have been limited to a single-word answer or selecting a phrase from a fixed set of phrases. These metrics limit the VidQA model…
Visual Semantic Role Labeling for Video Understanding
Arka Sadhu, Tanmay Gupta, Mark Yatskar +2
We propose a new framework for understanding and representing related salient events in a video using visual semantic role labeling. We represent videos as a set of related events,…
Utilizing Every Image Object for Semi-supervised Phrase Grounding
Haidong Zhu, Arka Sadhu, Zhaoheng Zheng +1
Phrase grounding models localize an object in the image given a referring expression. The annotated language queries available during training are limited, which also limits the va…
Video Object Grounding using Semantic Roles in Language Description
Arka Sadhu, Kan Chen, Ram Nevatia
We explore the task of Video Object Grounding (VOG), which grounds objects in videos referred to in natural language descriptions. Previous methods apply image grounding based algo…
Zero-Shot Grounding of Objects from Natural Language Queries
Arka Sadhu, Kan Chen, Ram Nevatia
A phrase grounding system localizes a particular object in an image referred to by a natural language query. In previous work, the phrases were restricted to have nouns that were e…