105 citations · 201 across the 9 of their papers we have counts for
28 papers
MixNorm: Test-Time Adaptation Through Online Normalization Estimation
Xuefeng Hu, Gokhan Uzunbas, Sirius Chen +4
We present a simple and effective way to estimate the batch-norm statistics during test time, to fast adapt a source model to target test samples. Known as Test-Time Adaptation, mo…
Visual Semantic Role Labeling for Video Understanding
Arka Sadhu, Tanmay Gupta, Mark Yatskar +2
We propose a new framework for understanding and representing related salient events in a video using visual semantic role labeling. We represent videos as a set of related events,…
SPAN: Spatial Pyramid Attention Network forImage Manipulation Localization
Xuefeng Hu, Zhihan Zhang, Zhenye Jiang +3
We present a novel framework, Spatial Pyramid Attention Network (SPAN) for detection and localization of multiple types of image manipulations. The proposed architecture efficientl…
CPARR: Category-based Proposal Analysis for Referring Relationships
Chuanzi He, Haidong Zhu, Jiyang Gao +2
The task of referring relationships is to localize subject and object entities in an image satisfying a relationship query, which is given in the form of \texttt{<subject, predicat…
Video Object Grounding using Semantic Roles in Language Description
Arka Sadhu, Kan Chen, Ram Nevatia
We explore the task of Video Object Grounding (VOG), which grounds objects in videos referred to in natural language descriptions. Previous methods apply image grounding based algo…
Zero-Shot Grounding of Objects from Natural Language Queries
Arka Sadhu, Kan Chen, Ram Nevatia
A phrase grounding system localizes a particular object in an image referred to by a natural language query. In previous work, the phrases were restricted to have nouns that were e…