60 citations · 98 across the 11 of their papers we have counts for
10 papers · 1 filter
TDViT: Temporal Dilated Video Transformer for Dense Video Tasks
Guanxiong Sun, Yang Hua, Guosheng Hu +1
Deep video models, for example, 3D CNNs or video transformers, have achieved promising performance on sparse video tasks, i.e., predicting one result per video. However, challenges…
Efficient One-stage Video Object Detection by Exploiting Temporal Consistency
Guanxiong Sun, Yang Hua, Guosheng Hu +1
Recently, one-stage detectors have achieved competitive accuracy and faster speed compared with traditional two-stage detectors on image data. However, in the field of video object…
MAMBA: Multi-level Aggregation via Memory Bank for Video Object Detection
Guanxiong Sun, Yang Hua, Guosheng Hu +1
State-of-the-art video object detection methods maintain a memory structure, either a sliding window or a memory queue, to enhance the current frame using attention mechanisms. How…
Imbalance Robust Softmax for Deep Embeeding Learning
Hao Zhu, Yang Yuan, Guosheng Hu +2
Deep embedding learning is expected to learn a metric space in which features have smaller maximal intra-class distance than minimal inter-class distance. In recent years, one rese…
Off-Policy Self-Critical Training for Transformer in Visual Paragraph Generation
Shiyang Yan, Yang Hua, Neil M. Robertson
Recently, several approaches have been proposed to solve language generation problems. Transformer is currently state-of-the-art seq-to-seq model in language generation. Reinforcem…
ParaCNN: Visual Paragraph Generation via Adversarial Twin Contextual CNNs
Shiyang Yan, Yang Hua, Neil Robertson
Image description generation plays an important role in many real-world applications, such as image retrieval, automatic navigation, and disabled people support. A well-developed t…