Associative Embedding: End-to-End Learning for Joint Detection and Grouping
arXiv:1611.05424
Abstract
We introduce associative embedding, a novel method for supervising convolutional neural networks for the task of detection and grouping. A number of computer vision problems can be framed in this manner including multi-person pose estimation, instance segmentation, and multi-object tracking. Usually the grouping of detections is achieved with multi-stage pipelines, instead we propose an approach that teaches a network to simultaneously output detections and group assignments. This technique can be easily integrated into any state-of-the-art network architecture that produces pixel-wise predictions. We show how to apply this method to both multi-person pose estimation and instance segmentation and report state-of-the-art performance for multi-person pose on the MPII and MS-COCO datasets.
Added results on MS-COCO and updated results on MPII
Cited by in corpus (169)
- Recent advances and clinical applications of deep learning in medical image analysis
- OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- YOLACT++: Better Real-time Instance Segmentation
- SOLOv2: Dynamic and Fast Instance Segmentation
- Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
- Deep High-Resolution Representation Learning for Visual Recognition
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- LCR-Net++: Multi-person 2D and 3D Pose Detection in Natural Images
- Deep Learning for Generic Object Detection: A Survey
- Semantic Instance Segmentation via Deep Metric Learning
- K-Net: Towards Unified Image Segmentation
- DeeperLab: Single-Shot Image Parser
- CornerNet-Lite: Efficient Keypoint Based Object Detection
- STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos
- CornerNet: Detecting Objects as Paired Keypoints
- RepPoints: Point Set Representation for Object Detection
- Simple Baselines for Human Pose Estimation and Tracking
- Rethinking on Multi-Stage Networks for Human Pose Estimation
- Bottom-up Object Detection by Grouping Extreme and Center Points
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- Deep Learning-Based Human Pose Estimation: A Survey
- DirectPose: Direct End-to-End Multi-Person Pose Estimation
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation
- 3D-BEVIS: Bird's-Eye-View Instance Segmentation
- Human Body Pose Estimation for Gait Identification: A Comprehensive Survey of Datasets and Models
- Deep High-Resolution Representation Learning for Human Pose Estimation
- Self-supervised Keypoint Correspondences for Multi-Person Pose Estimation and Tracking in Videos
- CenterNet3D: An Anchor Free Object Detector for Point Cloud
- Heatmap Regression via Randomized Rounding
- Object Discovery with a Copy-Pasting GAN
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- Single-shot 3D multi-person pose estimation in complex images
- SOLO: Segmenting Objects by Locations
- Learning Delicate Local Representations for Multi-Person Pose Estimation
- The Devil is in the Details: Delving into Unbiased Data Processing for Human Pose Estimation
- AP-10K: A Benchmark for Animal Pose Estimation in the Wild
- Bottom-Up Human Pose Estimation Via Disentangled Keypoint Regression
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- SPGNet: Semantic Prediction Guidance for Scene Parsing
- Associatively Segmenting Instances and Semantics in Point Clouds
- Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views
- Towards Real-Time Multi-Object Tracking
- Single-Stage Multi-Person Pose Machines
- Conditional Convolutions for Instance Segmentation
- Corner Proposal Network for Anchor-free, Two-stage Object Detection
- Rethinking the Heatmap Regression for Bottom-up Human Pose Estimation
- DanceIt: Music-inspired Dancing Video Synthesis
- Learning to Group: A Bottom-Up Framework for 3D Part Discovery in Unseen Categories
- TokenPose: Learning Keypoint Tokens for Human Pose Estimation
- Bottom-Up Human Pose Estimation by Ranking Heatmap-Guided Adaptive Keypoint Estimates
- Peeking into occluded joints: A novel framework for crowd pose estimation
- Instance Embedding Transfer to Unsupervised Video Object Segmentation
- Differentiable Hierarchical Graph Grouping for Multi-Person Pose Estimation
- Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation
- CenterFace: Joint Face Detection and Alignment Using Face as Point
- Multigrid Predictive Filter Flow for Unsupervised Learning on Videos
- KTAN: Knowledge Transfer Adversarial Network
- HRCenterNet: An Anchorless Approach to Chinese Character Segmentation in Historical Documents
- Graph Stacked Hourglass Networks for 3D Human Pose Estimation
- End-to-end Hand Mesh Recovery from a Monocular RGB Image
- Pose2Seg: Detection Free Human Instance Segmentation
- Single-Network Whole-Body Pose Estimation
- PolyTransform: Deep Polygon Transformer for Instance Segmentation
- Instance Segmentation with Point Supervision
- SG-Net: Spatial Granularity Network for One-Stage Video Instance Segmentation
- Pose Neural Fabrics Search
- VoxelPose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Graph-PCNN: Two Stage Human Pose Estimation with Graph Pose Refinement
- BANet: Bidirectional Aggregation Network with Occlusion Handling for Panoptic Segmentation
- AID: Pushing the Performance Boundary of Human Pose Estimation with Information Dropping Augmentation
- Improving Multi-Person Pose Estimation using Label Correction
- Generative Partition Networks for Multi-Person Pose Estimation
- PPGNet: Learning Point-Pair Graph for Line Segment Detection
- Explicit Shape Encoding for Real-Time Instance Segmentation
- Attentive Relational Networks for Mapping Images to Scene Graphs
- LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking
- Simple Pose: Rethinking and Improving a Bottom-up Approach for Multi-Person Pose Estimation
- PanoNet: Real-time Panoptic Segmentation through Position-Sensitive Feature Embedding
- Focus on Local: Detecting Lane Marker from Bottom Up via Key Point
- End-to-End Trainable Multi-Instance Pose Estimation with Transformers
- Monocular, One-stage, Regression of Multiple 3D People
- Embedding-based Instance Segmentation in Microscopy
- Learning to Refine Human Pose Estimation
- Multi-Person Pose Estimation with Enhanced Channel-wise and Spatial Information
- Spatial Feature Calibration and Temporal Fusion for Effective One-stage Video Instance Segmentation
- Semi-convolutional Operators for Instance Segmentation
- Towards High Performance Human Keypoint Detection
- Multi-person Articulated Tracking with Spatial and Temporal Embeddings
- When Human Pose Estimation Meets Robustness: Adversarial Algorithms and Benchmarks
- HMOR: Hierarchical Multi-Person Ordinal Relations for Monocular Multi-Person 3D Pose Estimation
- SCATTER: Selective Context Attentional Scene Text Recognizer
- Pose Recognition with Cascade Transformers
- SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation
- A Context-and-Spatial Aware Network for Multi-Person Pose Estimation
- HDNet: Human Depth Estimation for Multi-Person Camera-Space Localization
- ScaleNAS: One-Shot Learning of Scale-Aware Representations for Visual Recognition
- Differentiable Multi-Granularity Human Representation Learning for Instance-Aware Human Semantic Parsing
- Learning and Segmenting Dense Voxel Embeddings for 3D Neuron Reconstruction
- Instance Segmentation and Tracking with Cosine Embeddings and Recurrent Hourglass Networks
- SMAP: Single-Shot Multi-Person Absolute 3D Pose Estimation
- Combining detection and tracking for human pose estimation in videos
- Deep Instance Segmentation and Visual Servoing to Play Jenga with a Cost-Effective Robotic System
- Augmented Parallel-Pyramid Net for Attention Guided Pose-Estimation
- Instance and Panoptic Segmentation Using Conditional Convolutions
- Learning to Separate: Detecting Heavily-Occluded Objects in Urban Scenes
- Multi-scale Aggregation R-CNN for 2D Multi-person Pose Estimation
- 3DCFS: Fast and Robust Joint 3D Semantic-Instance Segmentation via Coupled Feature Selection
- Whole-Body Human Pose Estimation in the Wild
- TRB: A Novel Triplet Representation for Understanding 2D Human Body
- EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
- Spatial Shortcut Network for Human Pose Estimation
- Learning to Track Instances without Video Annotations
- Towards Fast and Accurate Multi-Person Pose Estimation on Mobile Devices
- A^2-FPN: Attention Aggregation based Feature Pyramid Network for Instance Segmentation
- Train Your Data Processor: Distribution-Aware and Error-Compensation Coordinate Decoding for Human Pose Estimation
- SimPose: Effectively Learning DensePose and Surface Normals of People from Simulated Data
- Turbo Learning Framework for Human-Object Interactions Recognition and Human Pose Estimation
- X-LineNet: Detecting Aircraft in Remote Sensing Images by a pair of Intersecting Line Segments
- Ordinal Depth Supervision for 3D Human Pose Estimation
- StarMap for Category-Agnostic Keypoint and Viewpoint Estimation
- FCPose: Fully Convolutional Multi-Person Pose Estimation with Dynamic Instance-Aware Convolutions
- PI-Net: Pose Interacting Network for Multi-Person Monocular 3D Pose Estimation
- InsPose: Instance-Aware Networks for Single-Stage Multi-Person Pose Estimation
- Real-Time Panoptic Segmentation from Dense Detections
- VolumeFusion: Deep Depth Fusion for 3D Scene Reconstruction
- Inception Convolution with Efficient Dilation Search
- Greedy Offset-Guided Keypoint Grouping for Human Pose Estimation
- Monocular 3D Multi-Person Pose Estimation by Integrating Top-Down and Bottom-Up Networks
- NADS-Net: A Nimble Architecture for Driver and Seat Belt Detection via Convolutional Neural Networks
- Orderly Dual-Teacher Knowledge Distillation for Lightweight Human Pose Estimation
- DeepACEv2: Automated Chromosome Enumeration in Metaphase Cell Images Using Deep Convolutional Neural Networks
- SpineOne: A One-Stage Detection Framework for Degenerative Discs and Vertebrae
- Attend to Who You Are: Supervising Self-Attention for Keypoint Detection and Instance-Aware Association
- Single-Image Piece-wise Planar 3D Reconstruction via Associative Embedding
- Multi-Frame Content Integration with a Spatio-Temporal Attention Mechanism for Person Video Motion Transfer
- Human-centric Relation Segmentation: Dataset and Solution
- Associative Embedding for Game-Agnostic Team Discrimination
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- Unifying Part Detection and Association for Recurrent Multi-Person Pose Estimation
- Single View Physical Distance Estimation using Human Pose
- SOLO: A Simple Framework for Instance Segmentation
- Semantic Segmentation and Object Detection Towards Instance Segmentation: Breast Tumor Identification
- PandaNet : Anchor-Based Single-Shot Multi-Person 3D Pose Estimation
- Pose estimator and tracker using temporal flow maps for limbs
- Efficient Human Pose Estimation with Depthwise Separable Convolution and Person Centroid Guided Joint Grouping
- Stereo Object Matching Network
- KPNet: Towards Minimal Face Detector
- Multi-Person Pose Estimation with Enhanced Feature Aggregation and Selection
- Explicit Spatiotemporal Joint Relation Learning for Tracking Human Pose
- FollowMeUp Sports: New Benchmark for 2D Human Keypoint Recognition
- Semantic Attention and Scale Complementary Network for Instance Segmentation in Remote Sensing Images
- GSTO: Gated Scale-Transfer Operation for Multi-Scale Feature Learning in Pixel Labeling
- Towards Good Practices for Multi-Person Pose Estimation
- A Top-down Approach to Articulated Human Pose Estimation and Tracking
- ICM-3D: Instantiated Category Modeling for 3D Instance Segmentation
- Self-Supervision and Spatial-Sequential Attention Based Loss for Multi-Person Pose Estimation
- Bottom-up Pose Estimation of Multiple Person with Bounding Box Constraint
- Bounding Box Embedding for Single Shot Person Instance Segmentation
- Multi-Level Network for High-Speed Multi-Person Pose Estimation
- MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual Network
- MaskPlus: Improving Mask Generation for Instance Segmentation
- VideoClick: Video Object Segmentation with a Single Click
- Pixel Consensus Voting for Panoptic Segmentation
- Indoor Panorama Planar 3D Reconstruction via Divide and Conquer
- Learning Spatial Context with Graph Neural Network for Multi-Person Pose Grouping
- Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation
- A Coarse-to-Fine Instance Segmentation Network with Learning Boundary Representation
- Iterative Greedy Matching for 3D Human Pose Tracking from Multiple Views