activity
20192021
collaborators

6 papers

cs.CV2021

Improving Object Detection and Attribute Recognition by Feature Entanglement Reduction

Zhaoheng Zheng, Arka Sadhu, Ram Nevatia

We explore object detection with two attributes: color and material. The task aims to simultaneously detect objects and infer their color and material. A straight-forward approach…

cs.CV2021

Video Question Answering with Phrases via Semantic Roles

Arka Sadhu, Kan Chen, Ram Nevatia

Video Question Answering (VidQA) evaluation metrics have been limited to a single-word answer or selecting a phrase from a fixed set of phrases. These metrics limit the VidQA model…

cs.CV2021

Visual Semantic Role Labeling for Video Understanding

Arka Sadhu, Tanmay Gupta, Mark Yatskar +2

We propose a new framework for understanding and representing related salient events in a video using visual semantic role labeling. We represent videos as a set of related events,…

cs.CV2020

Utilizing Every Image Object for Semi-supervised Phrase Grounding

Haidong Zhu, Arka Sadhu, Zhaoheng Zheng +1

Phrase grounding models localize an object in the image given a referring expression. The annotated language queries available during training are limited, which also limits the va…

cs.CV2020

Video Object Grounding using Semantic Roles in Language Description

Arka Sadhu, Kan Chen, Ram Nevatia

We explore the task of Video Object Grounding (VOG), which grounds objects in videos referred to in natural language descriptions. Previous methods apply image grounding based algo…

cs.CV2019

Zero-Shot Grounding of Objects from Natural Language Queries

Arka Sadhu, Kan Chen, Ram Nevatia

A phrase grounding system localizes a particular object in an image referred to by a natural language query. In previous work, the phrases were restricted to have nouns that were e…