Optimal Hyperparameters for Deep LSTM-Networks for Sequence Labeling Tasks
arXiv:1707.06799
Abstract
Selecting optimal parameters for a neural network architecture can often make the difference between mediocre and state-of-the-art performance. However, little is published which parameters and design choices should be evaluated or selected making the correct hyperparameter optimization often a "black art that requires expert experiences" (Snoek et al., 2012). In this paper, we evaluate the importance of different network design choices and hyperparameters for five common linguistic sequence tagging tasks (POS, Chunking, NER, Entity Recognition, and Event Detection). We evaluated over 50.000 different setups and found, that some parameters, like the pre-trained word embeddings or the last layer of the network, have a large impact on the performance, while other parameters, for example the number of LSTM layers or the number of recurrent units, are of minor importance. We give a recommendation on a configuration that performs well among different tasks.
34 pages. 9 page version of this paper published at EMNLP 2017
References in corpus (3)
Cited by in corpus (8)
- Multilingual Training and Cross-lingual Adaptation on CTC-based Acoustic Model
- Comparing Rule-based, Feature-based and Deep Neural Methods for De-identification of Dutch Medical Records
- Rethinking Generalization of Neural Models: A Named Entity Recognition Case Study
- Syllable-based Neural Named Entity Recognition for Myanmar Language
- Joint Learning of Word and Label Embeddings for Sequence Labelling in Spoken Language Understanding
- Multi-Task Learning for Argumentation Mining
- Quantization Loss Re-Learning Method
- Teaching a Machine to Diagnose a Heart Disease; Beginning from digitizing scanned ECGs to detecting the Brugada Syndrome (BrS)