One Model To Learn Them All
arXiv:1706.05137
Abstract
Deep learning yields great results across many fields, from speech recognition, image classification, to translation. But for each problem, getting a deep model to work well involves research into the architecture and a long period of tuning. We present a single model that yields good results on a number of problems spanning multiple domains. In particular, this single model is trained concurrently on ImageNet, multiple translation tasks, image captioning (COCO dataset), a speech recognition corpus, and an English parsing task. Our model architecture incorporates building blocks from multiple domains. It contains convolutional layers, an attention mechanism, and sparsely-gated layers. Each of these computational blocks is crucial for a subset of the tasks we train on. Interestingly, even if a block is not crucial for a task, we observe that adding it never hurts performance and in most cases improves it on all tasks. We also show that tasks with less data benefit largely from joint training with other tasks, while performance on large tasks degrades only slightly if at all.
References in corpus (3)
Cited by in corpus (35)
- Multi-Task Learning with Deep Neural Networks: A Survey
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Attention over Parameters for Dialogue Systems
- Image-Conditioned Graph Generation for Road Network Extraction
- Towards Universal Object Detection by Domain Attention
- Collaborative Intelligence: Challenges and Opportunities
- BADGER: Learning to (Learn [Learning Algorithms] through Multi-Agent Communication)
- Deep Learning for identifying radiogenomic associations in breast cancer
- Multitask Learning Strengthens Adversarial Robustness
- Machine Translation: A Literature Review
- Describe What to Change: A Text-guided Unsupervised Image-to-Image Translation Approach
- Bidirectional Scene Text Recognition with a Single Decoder
- TIMELY: Pushing Data Movements and Interfaces in PIM Accelerators Towards Local and in Time Domain
- SpeechNet: A Universal Modularized Model for Speech Processing Tasks
- Neural Skill Transfer from Supervised Language Tasks to Reading Comprehension
- Computation on Sparse Neural Networks: an Inspiration for Future Hardware
- HyperGrid: Efficient Multi-Task Transformers with Grid-wise Decomposable Hyper Projections
- Exploring Uncertainty in Conditional Multi-Modal Retrieval Systems
- Enhancing a Neurocognitive Shared Visuomotor Model for Object Identification, Localization, and Grasping With Learning From Auxiliary Tasks
- AutoSeM: Automatic Task Selection and Mixing in Multi-Task Learning
- Comprehensive and Efficient Data Labeling via Adaptive Model Scheduling
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge Distillation
- Multilingual Dialogue Generation with Shared-Private Memory
- Cross-lingual Data Transformation and Combination for Text Classification
- Software/Hardware Co-design for Multi-modal Multi-task Learning in Autonomous Systems
- Towards Robust Pattern Recognition: A Review
- A Dataset and Benchmarks for Multimedia Social Analysis
- Deep Artificial Intelligence for Fantasy Football Language Understanding
- Recurrent Stacking of Layers in Neural Networks: An Application to Neural Machine Translation
- Deep Unified Multimodal Embeddings for Understanding both Content and Users in Social Media Networks
- CARLS: Cross-platform Asynchronous Representation Learning System
- Formal Fields: A Framework to Automate Code Generation Across Domains
- One Network Fits All? Modular versus Monolithic Task Formulations in Neural Networks
- Interaction Networks: Using a Reinforcement Learner to train other Machine Learning algorithms
- Adaptive Precision Training for Resource Constrained Devices