Multilingual Language Processing From Bytes
arXiv:1512.00103
Abstract
We describe an LSTM-based model which we call Byte-to-Span (BTS) that reads text as bytes and outputs span annotations of the form [start, length, label] where start positions, lengths, and labels are separate entries in our vocabulary. Because we operate directly on unicode bytes rather than language-specific words or characters, we can analyze text in many languages with a single model. Due to the small vocabulary size, these multilingual models are very compact, but produce results similar to or better than the state-of- the-art in Part-of-Speech tagging and Named Entity Recognition that use only the provided training datasets (no external data sources). Our models are learning "from scratch" in that they do not rely on any elements of the standard pipeline in Natural Language Processing (including tokenization), and thus can run in standalone fashion on raw text.
References in corpus (8)
- Sequence to Sequence Learning with Neural Networks
- Improving neural networks by preventing co-adaptation of feature detectors
- Natural Language Processing (almost) from Scratch
- Recurrent Neural Network Regularization
- Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
- Text Understanding from Scratch
- Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
- Improved Transition-Based Parsing by Modeling Characters instead of Words with LSTMs
Cited by in corpus (16)
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Multi-Task Cross-Lingual Sequence Tagging from Scratch
- SyntaxNet Models for the CoNLL 2017 Shared Task
- Toward Mention Detection Robustness with Recurrent Neural Networks
- What to do about non-standard (or non-canonical) language in NLP
- Neural Morphological Tagging from Characters for Morphologically Rich Languages
- Chargrid: Towards Understanding 2D Documents
- Hierarchical Contextualized Representation for Named Entity Recognition
- Keystroke dynamics as signal for shallow syntactic parsing
- How Robust Are Character-Based Word Embeddings in Tagging and MT Against Wrod Scramlbing or Randdm Nouse?
- Complex Structure Leads to Overfitting: A Structure Regularization Decoding Method for Natural Language Processing
- Neural Architectures for Named Entity Recognition
- Towards A Multi-agent System for Online Hate Speech Detection
- Statistical Parametric Speech Synthesis Using Bottleneck Representation From Sequence Auto-encoder
- Sentiment Tagging with Partial Labels using Modular Architectures
- A Byte-sized Approach to Named Entity Recognition