330 citations · 387 across the 4 of their papers we have counts for
14 papers · 1 filter
Video Question Answering with Phrases via Semantic Roles
Arka Sadhu, Kan Chen, Ram Nevatia
Video Question Answering (VidQA) evaluation metrics have been limited to a single-word answer or selecting a phrase from a fixed set of phrases. These metrics limit the VidQA model…
Unbiased Teacher for Semi-Supervised Object Detection
Yen-Cheng Liu, Chih-Yao Ma, Zijian He +6
Semi-supervised learning, i.e., training networks with both labeled and unlabeled data, has made significant progress recently. However, existing works have primarily focused on im…
FBNetV3: Joint Architecture-Recipe Search using Predictor Pretraining
Xiaoliang Dai, Alvin Wan, Peizhao Zhang +8
Neural Architecture Search (NAS) yields state-of-the-art neural networks that outperform their best manually-designed counterparts. However, previous NAS methods search for archite…
CPARR: Category-based Proposal Analysis for Referring Relationships
Chuanzi He, Haidong Zhu, Jiyang Gao +2
The task of referring relationships is to localize subject and object entities in an image satisfying a relationship query, which is given in the form of \texttt{<subject, predicat…
FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions
Alvin Wan, Xiaoliang Dai, Peizhao Zhang +9
Differentiable Neural Architecture Search (DNAS) has demonstrated great success in designing state-of-the-art, efficient neural networks. However, DARTS-based DNAS's search space i…
Video Object Grounding using Semantic Roles in Language Description
Arka Sadhu, Kan Chen, Ram Nevatia
We explore the task of Video Object Grounding (VOG), which grounds objects in videos referred to in natural language descriptions. Previous methods apply image grounding based algo…