End-to-End Trainable Multi-Instance Pose Estimation with Transformers
arXiv:2103.12115
Abstract
We propose an end-to-end trainable approach for multi-instance pose estimation, called POET (POse Estimation Transformer). Combining a convolutional neural network with a transformer encoder-decoder architecture, we formulate multiinstance pose estimation from images as a direct set prediction problem. Our model is able to directly regress the pose of all individuals, utilizing a bipartite matching scheme. POET is trained using a novel set-based global loss that consists of a keypoint loss, a visibility loss and a class loss. POET reasons about the relations between multiple detected individuals and the full image context to directly predict their poses in parallel. We show that POET achieves high accuracy on the COCO keypoint detection task while having less parameters and higher inference speed than other bottom-up and top-down approaches. Moreover, we show successful transfer learning when applying POET to animal pose estimation. To the best of our knowledge, this model is the first end-to-end trainable multi-instance pose estimation method and we hope it will serve as a simple and promising alternative.
References in corpus (12)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Transformers in Vision: A Survey
- Language Models are Few-Shot Learners
- Improved Regularization of Convolutional Neural Networks with Cutout
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Zero-Shot Text-to-Image Generation
- Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
- A Primer on Motion Capture with Deep Learning: Principles, Pitfalls and Perspectives
- Deep Learning-Based Human Pose Estimation: A Survey
- Deep High-Resolution Representation Learning for Human Pose Estimation
- Generative Language Modeling for Automated Theorem Proving
- End-to-End Human Pose and Mesh Reconstruction with Transformers
Cited by in corpus (4)
- TokenPose: Learning Keypoint Tokens for Human Pose Estimation
- Test-Time Personalization with a Transformer for Human Pose Estimation
- Semi-Supervised 3D Hand-Object Poses Estimation with Interactions in Time
- Attend to Who You Are: Supervising Self-Attention for Keypoint Detection and Instance-Aware Association