Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
arXiv:1705.07115
Abstract
Numerous deep learning applications benefit from multi-task learning with multiple regression and classification objectives. In this paper we make the observation that the performance of such systems is strongly dependent on the relative weighting between each task's loss. Tuning these weights by hand is a difficult and expensive process, making multi-task learning prohibitive in practice. We propose a principled approach to multi-task deep learning which weighs multiple loss functions by considering the homoscedastic uncertainty of each task. This allows us to simultaneously learn various quantities with different units or scales in both classification and regression settings. We demonstrate our model learning per-pixel depth regression, semantic and instance segmentation from a monocular input image. Perhaps surprisingly, we show our model can learn multi-task weightings and outperform separate models trained individually on each task.
CVPR 2018
References in corpus (1)
Cited by in corpus (116)
- An Overview of Multi-Task Learning in Deep Neural Networks
- FairMOT: On the Fairness of Detection and Re-Identification in Multiple Object Tracking
- GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks
- Multi-Task Learning with Deep Neural Networks: A Survey
- DeeperLab: Single-Shot Image Parser
- Evaluating Bayesian Deep Learning Methods for Semantic Segmentation
- Libra R-CNN: Towards Balanced Learning for Object Detection
- Estimating Depth from RGB and Sparse Sensing
- Image coding for machines: an end-to-end learned approach
- Gated-SCNN: Gated Shape CNNs for Semantic Segmentation
- Scene Text Detection and Recognition: The Deep Learning Era
- Gradient Surgery for Multi-Task Learning
- DCN+: Mixed Objective and Deep Residual Coattention for Question Answering
- Faster Training of Mask R-CNN by Focusing on Instance Boundaries
- Predicting air quality via multimodal AI and satellite imagery
- Weakly-Supervised Video Object Grounding from Text by Loss Weighting and Object Interaction
- Weakly Supervised Learning of Instance Segmentation with Inter-pixel Relations
- Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation
- PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing
- Panoptic Feature Pyramid Networks
- Multi-Task Reinforcement Learning with Soft Modularization
- AdaTask: A Task-aware Adaptive Learning Rate Approach to Multi-task Learning
- DEFT: Detection Embeddings for Tracking
- Prediction of transonic flow over supercritical airfoils using geometric-encoding and deep-learning strategies
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- Reinforced Mnemonic Reader for Machine Reading Comprehension
- Bounding Box Regression with Uncertainty for Accurate Object Detection
- Stereo R-CNN based 3D Object Detection for Autonomous Driving
- DeepLab2: A TensorFlow Library for Deep Labeling
- Pyramidal Person Re-IDentification via Multi-Loss Dynamic Training
- Striking the Right Balance with Uncertainty
- Uncertainty in multitask learning: joint representations for probabilistic MR-only radiotherapy planning
- Hierarchical Attentive Recurrent Tracking
- SSAP: Single-Shot Instance Segmentation With Affinity Pyramid
- Scaling Wide Residual Networks for Panoptic Segmentation
- What Makes Training Multi-Modal Classification Networks Hard?
- Vehicle Attribute Recognition by Appearance: Computer Vision Methods for Vehicle Type, Make and Model Classification
- Deep Neural Networks for Marine Debris Detection in Sonar Images
- Multi-Channel Attention Selection GANs for Guided Image-to-Image Translation
- Multi-Channel Attention Selection GAN with Cascaded Semantic Guidance for Cross-View Image Translation
- AANet: Attribute Attention Network for Person Re-Identifications
- Probabilistic Deep Learning to Quantify Uncertainty in Air Quality Forecasting
- Uncertainty-aware Joint Salient Object and Camouflaged Object Detection
- End-to-End Multi-Task Learning with Attention
- Multimodal Motion Prediction with Stacked Transformers
- Temporal Attentive Alignment for Large-Scale Video Domain Adaptation
- Advancing Robust Underwater Acoustic Target Recognition through Multi-task Learning and Multi-Gate Mixture-of-Experts
- When Does Self-supervision Improve Few-shot Learning?
- Objects are Different: Flexible Monocular 3D Object Detection
- Multitask Learning for Fundamental Frequency Estimation in Music
- Self-tuning moving horizon estimation of nonlinear systems via physics-informed machine learning Koopman modeling
- The Edge of Depth: Explicit Constraints between Segmentation and Depth
- Reducing Uncertainty in Undersampled MRI Reconstruction with Active Acquisition
- The PAU Survey: Background light estimation with deep learning techniques
- Generalization in multitask deep neural classifiers: a statistical physics approach
- MTL-NAS: Task-Agnostic Neural Architecture Search towards General-Purpose Multi-Task Learning
- SceneCode: Monocular Dense Semantic Reconstruction using Learned Encoded Scene Representations
- Multi-Sensor 3D Object Box Refinement for Autonomous Driving
- Semi-convolutional Operators for Instance Segmentation
- Disentangled Image Matting
- Attending to Discriminative Certainty for Domain Adaptation
- Mix and match networks: encoder-decoder alignment for zero-pair image translation
- Multi-task neural networks by learned contextual inputs
- Uncertainty aware audiovisual activity recognition using deep Bayesian variational inference
- Multi-task Learning with Sample Re-weighting for Machine Reading Comprehension
- A Brief Review of Deep Multi-task Learning and Auxiliary Task Learning
- Auxiliary Tasks Speed Up Learning PointGoal Navigation
- Machine-learned 3D Building Vectorization from Satellite Imagery
- Rethinking Class Relations: Absolute-relative Supervised and Unsupervised Few-shot Learning
- Driven to Distraction: Self-Supervised Distractor Learning for Robust Monocular Visual Odometry in Urban Environments
- Uncertainty-Aware Deep Calibrated Salient Object Detection
- Differentiable Multi-Granularity Human Representation Learning for Instance-Aware Human Semantic Parsing
- Semi-supervised Skin Detection by Network with Mutual Guidance
- BAR: Bayesian Activity Recognition using variational inference
- Deep Multicameral Decoding for Localizing Unoccluded Object Instances from a Single RGB Image
- AutoSeM: Automatic Task Selection and Mixing in Multi-Task Learning
- DF-Net: Unsupervised Joint Learning of Depth and Flow using Cross-Task Consistency
- Boundary Aware U-Net for Glacier Segmentation
- Joint Pose and Shape Estimation of Vehicles from LiDAR Data
- A Decoupled Uncertainty Model for MRI Segmentation Quality Estimation
- Exploring Uncertainty in Conditional Multi-Modal Retrieval Systems
- Uncertainty-Aware Physically-Guided Proxy Tasks for Unseen Domain Face Anti-spoofing
- Instance-Level Task Parameters: A Robust Multi-task Weighting Framework
- A Modulation Layer to Increase Neural Network Robustness Against Data Quality Issues
- Left Ventricle Segmentation and Quantification from Cardiac Cine MR Images via Multi-task Learning
- One-Vote Veto: Semi-Supervised Learning for Low-Shot Glaucoma Diagnosis
- Your "Flamingo" is My "Bird": Fine-Grained, or Not
- Auto-Retoucher(ART) - A framework for Background Replacement and Image Editing
- Renovating Parsing R-CNN for Accurate Multiple Human Parsing
- Cerberus Transformer: Joint Semantic, Affordance and Attribute Parsing
- Learning Probabilistic Ordinal Embeddings for Uncertainty-Aware Regression
- Learning to Abstract and Predict Human Actions
- SCNet: A Generalized Attention-based Model for Crack Fault Segmentation
- SUB-Depth: Self-distillation and Uncertainty Boosting Self-supervised Monocular Depth Estimation
- Interleaved Multitask Learning with Energy Modulated Learning Progress
- Monocular Outdoor Semantic Mapping with a Multi-task Network
- Exceeding the Limits of Visual-Linguistic Multi-Task Learning
- Graph Convolution for Multimodal Information Extraction from Visually Rich Documents
- Learning a Unified Embedding for Visual Search at Pinterest
- Joint Spatial-Temporal Optimization for Stereo 3D Object Tracking
- Neural Multi-Task Learning for Teacher Question Detection in Online Classrooms
- Deep Likelihood Network for Image Restoration with Multiple Degradation Levels
- Eliminating the Blind Spot: Adapting 3D Object Detection and Monocular Depth Estimation to 360° Panoramic Imagery
- Wrapped Loss Function for Regularizing Nonconforming Residual Distributions
- Transfer Learning in Visual and Relational Reasoning
- Efficient Multi-Domain Network Learning by Covariance Normalization
- Associative Embedding for Game-Agnostic Team Discrimination
- Incremental Scene Synthesis
- Learning to Generalize One Sample at a Time with Self-Supervision
- Dual Projection Generative Adversarial Networks for Conditional Image Generation
- Metric-based Regularization and Temporal Ensemble for Multi-task Learning using Heterogeneous Unsupervised Tasks
- Visually Aware Skip-Gram for Image Based Recommendations
- RGB-based Semantic Segmentation Using Self-Supervised Depth Pre-Training
- Multiple Classification with Split Learning
- W-net: Simultaneous segmentation of multi-anatomical retinal structures using a multi-task deep neural network
- High Diversity Attribute Guided Face Generation with GANs