67 citations · 609 across the 47 of their papers we have counts for
71 papers
Towards Lightweight Transformer via Group-wise Transformation for Vision-and-Language Tasks
Gen Luo, Yiyi Zhou, Xiaoshuai Sun +5
Despite the exciting performance, Transformer is criticized for its excessive parameters and computation cost. However, compressing Transformer remains as an open problem due to it…
ASFD: Automatic and Scalable Face Detector
Jian Li, Bin Zhang, Yabiao Wang +6
Along with current multi-scale based detectors, Feature Aggregation and Enhancement (FAE) modules have shown superior performance gains for cutting-edge object detection. However,…
LSTC: Boosting Atomic Action Detection with Long-Short-Term Context
Yuxi Li, Boshen Zhang, Jian Li +5
In this paper, we place the atomic action detection problem into a Long-Short Term Context (LSTC) to analyze how the temporal reliance among video signals affect the action detecti…
Transformer-based Dual Relation Graph for Multi-label Image Recognition
Jiawei Zhao, Ke Yan, Yifan Zhao +3
The simultaneous recognition of multiple objects in one image remains a challenging task, spanning multiple events in the recognition field such as various object scales, inconsist…
Spatiotemporal Inconsistency Learning for DeepFake Video Detection
Zhihao Gu, Yang Chen, Taiping Yao +4
The rapid development of facial manipulation techniques has aroused public concerns in recent years. Following the success of deep learning, existing methods always formulate DeepF…
Distributed Attention for Grounded Image Captioning
Nenglun Chen, Xingjia Pan, Runnan Chen +7
We study the problem of weakly supervised grounded image captioning. That is, given an image, the goal is to automatically generate a sentence describing the context of the image w…