10 citations · 20 across the 5 of their papers we have counts for
5 papers
Video Event Extraction via Tracking Visual States of Arguments
Guang Yang, Manling Li, Jiajie Zhang +3
Video event extraction aims to detect salient events from a video and identify the arguments for each event as well as their semantic roles. Existing methods focus on capturing the…
Video in 10 Bits: Few-Bit VideoQA for Efficiency and Privacy
Shiyuan Huang, Robinson Piramuthu, Shih-Fu Chang +1
In Video Question Answering (VideoQA), answering general questions about a video requires its visual information. Yet, video often contains redundant information irrelevant to the…
Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks
Zhecan Wang, Noel Codella, Yen-Chun Chen +8
Cross-modal encoders for vision-language (VL) tasks are often pretrained with carefully curated vision-language datasets. While these datasets reach an order of 10 million samples,…
Fine-Grained Visual Entailment
Christopher Thomas, Yipeng Zhang, Shih-Fu Chang
Visual entailment is a recently proposed multimodal reasoning task where the goal is to predict the logical relationship of a piece of text to an image. In this paper, we propose a…
Learning Visual Commonsense for Robust Scene Graph Generation
Alireza Zareian, Zhecan Wang, Haoxuan You +1
Scene graph generation models understand the scene through object and predicate recognition, but are prone to mistakes due to the challenges of perception in the wild. Perception e…