On the Relevance of Temporal Features for Medical Ultrasound Video Recognition
arXiv:2310.10453 · doi:10.1007/978-3-031-43895-0_70
Abstract
Many medical ultrasound video recognition tasks involve identifying key anatomical features regardless of when they appear in the video suggesting that modeling such tasks may not benefit from temporal features. Correspondingly, model architectures that exclude temporal features may have better sample efficiency. We propose a novel multi-head attention architecture that incorporates these hypotheses as inductive priors to achieve better sample efficiency on common ultrasound tasks. We compare the performance of our architecture to an efficient 3D CNN video recognition model in two settings: one where we expect not to require temporal features and one where we do. In the former setting, our model outperforms the 3D CNN - especially when we artificially limit the training data. In the latter, the outcome reverses. These results suggest that expressive time-independent models may be more effective than state-of-the-art video recognition models for some common ultrasound tasks in the low-data regime.
14 pages, 4 figures, published in MICCAI 23
References in corpus (4)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition
- Spatial Temporal Transformer Network for Skeleton-based Action Recognition
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights