most citedLearning Visual Commonsense for Robust Scene Graph Generation

10 citations · 20 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20222 cited

Video Event Extraction via Tracking Visual States of Arguments

Guang Yang, Manling Li, Jiajie Zhang +3

Video event extraction aims to detect salient events from a video and identify the arguments for each event as well as their semantic roles. Existing methods focus on capturing the…

cs.CV2022

Video in 10 Bits: Few-Bit VideoQA for Efficiency and Privacy

Shiyuan Huang, Robinson Piramuthu, Shih-Fu Chang +1

In Video Question Answering (VideoQA), answering general questions about a video requires its visual information. Yet, video often contains redundant information irrelevant to the…

cs.CV20228 cited

Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks

Zhecan Wang, Noel Codella, Yen-Chun Chen +8

Cross-modal encoders for vision-language (VL) tasks are often pretrained with carefully curated vision-language datasets. While these datasets reach an order of 10 million samples,…

cs.CV2022

Fine-Grained Visual Entailment

Christopher Thomas, Yipeng Zhang, Shih-Fu Chang

Visual entailment is a recently proposed multimodal reasoning task where the goal is to predict the logical relationship of a piece of text to an image. In this paper, we propose a…

cs.CV202010 cited

Learning Visual Commonsense for Robust Scene Graph Generation

Alireza Zareian, Zhecan Wang, Haoxuan You +1

Scene graph generation models understand the scene through object and predicate recognition, but are prone to mistakes due to the challenges of perception in the wild. Perception e…