Multi-Task Learning as Multi-Objective Optimization
arXiv:1810.04650
Abstract
In multi-task learning, multiple tasks are solved jointly, sharing inductive bias between them. Multi-task learning is inherently a multi-objective problem because different tasks may conflict, necessitating a trade-off. A common compromise is to optimize a proxy objective that minimizes a weighted linear combination of per-task losses. However, this workaround is only valid when the tasks do not compete, which is rarely the case. In this paper, we explicitly cast multi-task learning as multi-objective optimization, with the overall objective of finding a Pareto optimal solution. To this end, we use algorithms developed in the gradient-based multi-objective optimization literature. These algorithms are not directly applicable to large-scale learning problems since they scale poorly with the dimensionality of the gradients and the number of tasks. We therefore propose an upper bound for the multi-objective loss and show that it can be optimized efficiently. We further prove that optimizing this upper bound yields a Pareto optimal solution under realistic assumptions. We apply our method to a variety of multi-task deep learning problems including digit classification, scene understanding (joint semantic segmentation, instance segmentation, and depth estimation), and multi-label classification. Our method produces higher-performing models than recent multi-task learning formulations or per-task training.
In Neural Information Processing Systems (NeurIPS) 2018
Cited by in corpus (113)
- FairMOT: On the Fairness of Detection and Re-Identification in Multiple Object Tracking
- Multi-Task Learning for Dense Prediction Tasks: A Survey
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Multi-Task Learning with Deep Neural Networks: A Survey
- Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection
- Physics-informed neural networks for the shallow-water equations on the sphere
- Blind Backdoors in Deep Learning Models
- Gradient Surgery for Multi-Task Learning
- Multi-Objective Matrix Normalization for Fine-grained Visual Recognition
- Adapting Auxiliary Losses Using Gradient Similarity
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
- Multi-Task Reinforcement Learning with Soft Modularization
- AdaTask: A Task-aware Adaptive Learning Rate Approach to Multi-task Learning
- Context-Aware Embeddings for Automatic Art Analysis
- Self-Regulated Learning for Egocentric Video Activity Anticipation
- Motif-based Graph Self-Supervised Learning for Molecular Property Prediction
- Learning to Branch for Multi-Task Learning
- MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale
- Pareto-Optimal Bit Allocation for Collaborative Intelligence
- Towards Real-Time Multi-Object Tracking
- Regularizing Deep Multi-Task Networks using Orthogonal Gradients
- Learning the Pareto Front with Hypernetworks
- Efficiently Identifying Task Groupings for Multi-Task Learning
- Improved Schemes for Episodic Memory-based Lifelong Learning
- Multi-Gradient Descent for Multi-Objective Recommender Systems
- Controllable Pareto Multi-Task Learning
- Multitask Learning for Scalable and Dense Multilayer Bayesian Map Inference
- H3DNet: 3D Object Detection Using Hybrid Geometric Primitives
- SAMBA: Safe Model-Based & Active Reinforcement Learning
- One Model to Serve All: Star Topology Adaptive Recommender for Multi-Domain CTR Prediction
- When Does Self-supervision Improve Few-shot Learning?
- Multi-task learning for virtual flow metering
- Auxiliary Task Reweighting for Minimum-data Learning
- Machine Learning-Assisted Thermoelectric Cooling for On-Demand Multi-Hotspot Thermal Management
- Code-switched inspired losses for generic spoken dialog representations
- Generalization in multitask deep neural classifiers: a statistical physics approach
- MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps
- Decomposition Multi-Objective Evolutionary Optimization: From State-of-the-Art to Future Opportunities
- Secure Watermark for Deep Neural Networks with Multi-task Learning
- An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
- PROUD: PaRetO-gUided Diffusion Model for Multi-objective Generation
- Multi-Objective Meta Learning
- Multi-script Handwritten Digit Recognition Using Multi-task Learning
- A Simple General Approach to Balance Task Difficulty in Multi-Task Learning
- Measuring and Harnessing Transference in Multi-Task Learning
- Image Enhanced Rotation Prediction for Self-Supervised Learning
- Parallax Attention for Unsupervised Stereo Correspondence Learning
- Auxiliary Learning for Deep Multi-task Learning
- Small Towers Make Big Differences
- The Fairness-Accuracy Pareto Front
- Multi-Domain Multi-Task Rehearsal for Lifelong Learning
- Embedding Adaptation is Still Needed for Few-Shot Learning
- Speech Representation Learning Through Self-supervised Pretraining And Multi-task Finetuning
- Boosting Supervision with Self-Supervision for Few-shot Learning
- Knowledge Distillation for Multi-task Learning
- Auxiliary Task Update Decomposition: The Good, The Bad and The Neutral
- Multi-Task Variational Information Bottleneck
- Scalable Multitask Learning Using Gradient-based Estimation of Task Affinity
- Reinforcement Learning for Joint Optimization of Multiple Rewards
- Cross-Task Consistency Learning Framework for Multi-Task Learning
- Auxiliary Learning by Implicit Differentiation
- Multi-Objective Neural Architecture Search Based on Diverse Structures and Adaptive Recommendation
- Avoiding hashing and encouraging visual semantics in referential emergent language games
- Multitask Learning with Single Gradient Step Update for Task Balancing
- Learning Multiple Dense Prediction Tasks from Partially Annotated Data
- Multi-Objective Learning to Predict Pareto Fronts Using Hypervolume Maximization
- Learning Boost by Exploiting the Auxiliary Task in Multi-task Domain
- Multi-Task Adversarial Attack
- Modeling Online Behavior in Recommender Systems: The Importance of Temporal Context
- Follow the bisector: a simple method for multi-objective optimization
- RotoGrad: Gradient Homogenization in Multitask Learning
- Relationship Explainable Multi-objective Reinforcement Learning with Semantic Explainability Generation
- Instance-Level Task Parameters: A Robust Multi-task Weighting Framework
- Pareto Self-Supervised Training for Few-Shot Learning
- Multi-Task Multicriteria Hyperparameter Optimization
- Multi-task Learning by Leveraging the Semantic Information
- Learning Invariant Representations across Domains and Tasks
- Multi-task problems are not multi-objective
- Self-Evolutionary Optimization for Pareto Front Learning
- PolyViT: Co-training Vision Transformers on Images, Videos and Audio
- Momentum-based Gradient Methods in Multi-Objective Recommendation
- Image interpretation by iterative bottom-up top-down processing
- Multi-task Supervised Learning via Cross-learning
- Spinning Sequence-to-Sequence Models with Meta-Backdoors
- Sequential Reptile: Inter-Task Gradient Alignment for Multilingual Learning
- Scalable Unidirectional Pareto Optimality for Multi-Task Learning with Constraints
- Multi-path Neural Networks for On-device Multi-domain Visual Classification
- Boosting Supervised Learning Performance with Co-training
- Pareto-Optimal Allocation of Transactive Energy at Market Equilibrium in Distribution Systems: A Constrained Vector Optimization Approach
- Interleaved Multitask Learning with Energy Modulated Learning Progress
- Generative Modeling for Multi-task Visual Learning
- Adaptive Distillation: Aggregating Knowledge from Multiple Paths for Efficient Distillation
- HydaLearn: Highly Dynamic Task Weighting for Multi-task Learning with Auxiliary Tasks
- Blind Image Super-Resolution with Spatial Context Hallucination
- Multi-Task Learning of Query Intent and Named Entities using Transfer Learning
- Real-time 3D Object Detection using Feature Map Flow
- Principal Gradient Direction and Confidence Reservoir Sampling for Continual Learning
- Fast Line Search for Multi-Task Learning
- Balancing Average and Worst-case Accuracy in Multitask Learning
- Learning Parallax Attention for Stereo Image Super-Resolution
- Disentangling Transfer and Interference in Multi-Domain Learning
- CACTUS: Detecting and Resolving Conflicts in Objective Functions
- A new approach to forecast service parts demand by integrating user preferences into multi-objective optimization
- A Novel Perspective to Zero-shot Learning: Towards an Alignment of Manifold Structures via Semantic Feature Expansion
- Unlocking the Full Potential of Small Data with Diverse Supervision
- PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data
- Parity Partition Coding for Sharp Multi-Label Classification
- Multiple Classification with Split Learning
- Deep Learning for CSI Feedback Based on Superimposed Coding
- Intra-Model Collaborative Learning of Neural Networks
- Fully differentiable model discovery
- Adjacency List Oriented Relational Fact Extraction via Adaptive Multi-task Learning
- Safety Aware Reinforcement Learning (SARL)