8 citations · 31 across the 10 of their papers we have counts for
7 papers
Detecting and Grounding Multi-Modal Media Manipulation and Beyond
Rui Shao, Tianxing Wu, Jianlong Wu +2
Misinformation has become a pressing issue. Fake media, in both visual and textual forms, is widespread on the web. While various deepfake detection and text fake news detection me…
Temporal Sentence Grounding in Streaming Videos
Tian Gan, Xiao Wang, Yan Sun +3
This paper aims to tackle a novel task - Temporal Sentence Grounding in Streaming Videos (TSGSV). The goal of TSGSV is to evaluate the relevance between a video stream and a given…
OFAR: A Multimodal Evidence Retrieval Framework for Illegal Live-streaming Identification
Lin Dengtian, Ma Yang, Li Yuhong +3
Illegal live-streaming identification, which aims to help live-streaming platforms immediately recognize the illegal behaviors in the live-streaming, such as selling precious and e…
Micro-video Tagging via Jointly Modeling Social Influence and Tag Relation
Xiao Wang, Tian Gan, Yinwei Wei +3
The last decade has witnessed the proliferation of micro-videos on various user-generated content platforms. According to our statistics, around 85.7\% of micro-videos lack annotat…
Visual Perturbation-aware Collaborative Learning for Overcoming the Language Prior Problem
Yudong Han, Liqiang Nie, Jianhua Yin +2
Several studies have recently pointed that existing Visual Question Answering (VQA) models heavily suffer from the language prior problem, which refers to capturing superficial sta…
Semantic-aware Modular Capsule Routing for Visual Question Answering
Yudong Han, Jianhua Yin, Jianlong Wu +2
Visual Question Answering (VQA) is fundamentally compositional in nature, and many questions are simply answered by decomposing them into modular sub-problems. The recent proposed…