Local-Global Context Aware Transformer for Language-Guided Video Segmentation
arXiv:2203.09773 · doi:10.1109/TPAMI.2023.3262578
Abstract
We explore the task of language-guided video segmentation (LVS). Previous algorithms mostly adopt 3D CNNs to learn video representation, struggling to capture long-term context and easily suffering from visual-linguistic misalignment. In light of this, we present Locater (local-global context aware Transformer), which augments the Transformer architecture with a finite memory so as to query the entire video with the language expression in an efficient manner. The memory is designed to involve two components -- one for persistently preserving global video content, and one for dynamically gathering local temporal context and segmentation history. Based on the memorized local-global context and the particular content of each frame, Locater holistically and flexibly comprehends the expression as an adaptive query vector for each frame. The vector is used to query the corresponding frame for mask generation. The memory also allows Locater to process videos with linear time complexity and constant size memory, while Transformer-style self-attention computation scales quadratically with sequence length. To thoroughly examine the visual grounding capability of LVS models, we contribute a new LVS dataset, A2D-S+, which is built upon A2D-S dataset but poses increased challenges in disambiguating among similar objects. Experiments on three LVS datasets and our A2D-S+ show that Locater outperforms previous state-of-the-arts. Further, we won the 1st place in the Referring Video Object Segmentation Track of the 3rd Large-scale Video Object Segmentation Challenge, where Locater served as the foundation for the winning solution. Our code and dataset are available at: https://github.com/leonnnop/Locater
Accepted by TPAMI. Code, data: https://github.com/leonnnop/Locater
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Sequence to Sequence Learning with Neural Networks
- Longformer: The Long-Document Transformer
- Linformer: Self-Attention with Linear Complexity
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers
- Referring Transformer: A One-step Approach to Multi-task Visual Grounding
- GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models
- Decoupling Features in Hierarchical Propagation for Video Object Segmentation
- Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation
- Towards Data-and Knowledge-Driven Artificial Intelligence: A Survey on Neuro-Symbolic Computing
Cited by in corpus (4)
- Efficient Long-Short Temporal Attention Network for Unsupervised Video Object Segmentation
- Deep learning-based blind image super-resolution with iterative kernel reconstruction and noise estimation
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- Forward Consistency Learning with Gated Context Aggregation for Video Anomaly Detection