Scaling up sign spotting through sign language dictionaries
arXiv:2205.04152 · doi:10.1007/s11263-022-01589-6
Abstract
The focus of this work is - given a video of an isolated sign, our task is to identify and it has been signed in a continuous, co-articulated sign language video. To achieve this sign spotting task, we train a model using multiple types of available supervision by: (1) existing footage which is sparsely labelled using mouthing cues; (2) associated subtitles (readily available translations of the signed content) which provide additional ; (3) words (for which no co-articulated labelled examples are available) in visual sign language dictionaries to enable novel sign spotting. These three tasks are integrated into a unified learning framework using the principles of Noise Contrastive Estimation and Multiple Instance Learning. We validate the effectiveness of our approach on low-shot sign spotting benchmarks. In addition, we contribute a machine-readable British Sign Language (BSL) dictionary dataset of isolated signs, BSLDict, to facilitate study of this task. The dataset, models and code are available at our project page.
Appears in: 2022 International Journal of Computer Vision (IJCV). 25 pages. arXiv admin note: substantial text overlap with arXiv:2010.04002
References in corpus (7)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Few-Shot Adversarial Domain Adaptation
- Towards Zero-shot Sign Language Recognition
- Seeing wake words: Audio-visual Keyword Spotting
- BBC-Oxford British Sign Language Dataset
- Signs in time: Encoding human motion as a temporal image
- Zero-Shot Sign Language Recognition: Can Textual Data Uncover Sign Languages?