End-to-end Temporal Action Detection with Transformer
arXiv:2106.10271 · doi:10.1109/TIP.2022.3195321
Abstract
Temporal action detection (TAD) aims to determine the semantic label and the temporal interval of every action instance in an untrimmed video. It is a fundamental and challenging task in video understanding. Previous methods tackle this task with complicated pipelines. They often need to train multiple networks and involve hand-designed operations, such as non-maximal suppression and anchor generation, which limit the flexibility and prevent end-to-end learning. In this paper, we propose an end-to-end Transformer-based method for TAD, termed TadTR. Given a small set of learnable embeddings called action queries, TadTR adaptively extracts temporal context information from the video for each query and directly predicts action instances with the context. To adapt Transformer to TAD, we propose three improvements to enhance its locality awareness. The core is a temporal deformable attention module that selectively attends to a sparse set of key snippets in a video. A segment refinement mechanism and an actionness regression head are designed to refine the boundaries and confidence of the predicted instances, respectively. With such a simple pipeline, TadTR requires lower computation cost than previous detectors, while preserving remarkable performance. As a self-contained detector, it achieves state-of-the-art performance on THUMOS14 (56.7% mAP) and HACS Segments (32.09% mAP). Combined with an extra action classifier, it obtains 36.75% mAP on ActivityNet-1.3. Code is available at https://github.com/xlliu7/TadTR.
Accepted by IEEE Transactions on Image Processing (TIP). Code: https://github.com/xlliu7/TadTR
References in corpus (9)
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Is Space-Time Attention All You Need for Video Understanding?
- Single Shot Temporal Action Detection
- TransCrowd: weakly-supervised crowd counting with transformers
- Revisiting Anchor Mechanisms for Temporal Action Localization
- CUHK & ETHZ & SIAT Submission to ActivityNet Challenge 2016
- ActivityNet Challenge 2017 Summary
- Activity Graph Transformer for Temporal Action Localization
- Temporal Action Proposal Generation with Transformers
Cited by in corpus (9)
- Hyperspectral Image Denoising via Spatial-Spectral Recurrent Transformer
- A Semantic and Motion-Aware Spatiotemporal Transformer Network for Action Detection
- Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
- PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference Transformer
- Low-power, Continuous Remote Behavioral Localization with Event Cameras
- UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
- Towards High-Quality Temporal Action Detection with Sparse Proposals
- SPAN: Continuous Modeling of Suspicion Progression for Temporal Intention Localization
- ArthroPhase: A Novel Dataset and Method for Phase Recognition in Arthroscopic Video