Character-Aware Neural Language Models
arXiv:1508.06615
Abstract
We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway network over characters, whose output is given to a long short-term memory (LSTM) recurrent neural network language model (RNN-LM). On the English Penn Treebank the model is on par with the existing state-of-the-art despite having 60% fewer parameters. On languages with rich morphology (Arabic, Czech, French, German, Spanish, Russian), the model outperforms word-level/morpheme-level LSTM baselines, again with fewer parameters. The results suggest that on many languages, character inputs are sufficient for language modeling. Analysis of word representations obtained from the character composition part of the model reveals that the model is able to encode, from characters only, both semantic and orthographic information.
AAAI 2016
References in corpus (10)
- Sequence to Sequence Learning with Neural Networks
- Improving neural networks by preventing co-adaptation of feature detectors
- Natural Language Processing (almost) from Scratch
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Recurrent Neural Network Regularization
- Training Very Deep Networks
- Compositional Morphology for Word Representations and Language Modelling
- Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
- CNN: A Convolutional Architecture for Word Sequence Prediction
- Probabilistic Modelling of Morphologically Rich Languages
Cited by in corpus (259)
- Pre-trained Models for Natural Language Processing: A Survey
- Recent Trends in Deep Learning Based Natural Language Processing
- Federated Learning for Mobile Keyboard Prediction
- A Survey on Recent Advances in Named Entity Recognition from Deep Learning models
- Born Again Neural Networks
- Enriching Word Vectors with Subword Information
- Deep Learning Based Text Classification: A Comprehensive Review
- Recent Advances in Convolutional Neural Networks
- What do Neural Machine Translation Models Learn about Morphology?
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- Arabic natural language processing: An overview
- On Adversarial Examples for Character-Level Neural Machine Translation
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Glyce: Glyph-vectors for Chinese Character Representations
- Big Code != Big Vocabulary: Open-Vocabulary Models for Source Code
- CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation
- Neural Summarization by Extracting Sentences and Words
- A Deep Network Model for Paraphrase Detection in Short Text Messages
- DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
- Supervised and Unsupervised Neural Approaches to Text Readability
- Sequence-Level Knowledge Distillation
- Discourse-Based Objectives for Fast Unsupervised Sentence Representation Learning
- ChemGAN challenge for drug discovery: can AI reproduce natural chemical diversity?
- A Survey of Deep Learning Techniques for Neural Machine Translation
- Measuring Depression Symptom Severity from Spoken Language and 3D Facial Expressions
- End-to-End ASR-free Keyword Search from Speech
- R-Transformer: Recurrent Neural Network Enhanced Transformer
- Character-level and Multi-channel Convolutional Neural Networks for Large-scale Authorship Attribution
- Deep Enhanced Representation for Implicit Discourse Relation Recognition
- A Unified Model for Opinion Target Extraction and Target Sentiment Prediction
- Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
- Semi-supervised Word Sense Disambiguation with Neural Models
- Ordered Neurons: Integrating Tree Structures into Recurrent Neural Networks
- Do Convolutional Networks need to be Deep for Text Classification ?
- BanFakeNews: A Dataset for Detecting Fake News in Bangla
- Distance-based Self-Attention Network for Natural Language Inference
- Neural Extractive Summarization with Side Information
- Attending to Characters in Neural Sequence Labeling Models
- GraphIE: A Graph-Based Framework for Information Extraction
- Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
- Dynamic Evaluation of Neural Sequence Models
- Adaptive Input Representations for Neural Language Modeling
- Dual Rectified Linear Units (DReLUs): A Replacement for Tanh Activation Functions in Quasi-Recurrent Neural Networks
- Universal Neural Machine Translation for Extremely Low Resource Languages
- Words Can Shift: Dynamically Adjusting Word Representations Using Nonverbal Behaviors
- Words or Characters? Fine-grained Gating for Reading Comprehension
- Compression of Recurrent Neural Networks for Efficient Language Modeling
- LightRNN: Memory and Computation-Efficient Recurrent Neural Networks
- Clinical Concept Extraction with Contextual Word Embedding
- Neural ranking models for document retrieval
- RNNs as psycholinguistic subjects: Syntactic state and grammatical dependency
- Multi-task Prediction of Disease Onsets from Longitudinal Lab Tests
- Alternative structures for character-level RNNs
- Faster On-Device Training Using New Federated Momentum Algorithm
- Deep Independently Recurrent Neural Network (IndRNN)
- Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech
- An Updated Duet Model for Passage Re-ranking
- Enhancing lexical-based approach with external knowledge for Vietnamese multiple-choice machine reading comprehension
- A Survey of Orthographic Information in Machine Translation
- Probabilistic FastText for Multi-Sense Word Embeddings
- Montage: A Neural Network Language Model-Guided JavaScript Engine Fuzzer
- Compressing Word Embeddings via Deep Compositional Code Learning
- Papaya: Practical, Private, and Scalable Federated Learning
- Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension
- CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters
- Comparing Neural- and N-Gram-Based Language Models for Word Segmentation
- Who Needs Words? Lexicon-Free Speech Recognition
- Enhance word representation for out-of-vocabulary on Ubuntu dialogue corpus
- Towards a Robust Deep Neural Network in Texts: A Survey
- Modeling Vocabulary for Big Code Machine Learning
- Spelling Error Correction Using a Nested RNN Model and Pseudo Training Data
- Evaluating Text GANs as Language Models
- 3D Gated Recurrent Fusion for Semantic Scene Completion
- Neural Language Modeling by Jointly Learning Syntax and Lexicon
- Subword-augmented Embedding for Cloze Reading Comprehension
- Hierarchical Character Embeddings: Learning Phonological and Semantic Representations in Languages of Logographic Origin using Recursive Neural Networks
- 75 Languages, 1 Model: Parsing Universal Dependencies Universally
- Is Word Segmentation Necessary for Deep Learning of Chinese Representations?
- QA4IE: A Question Answering based Framework for Information Extraction
- Binary Black-box Evasion Attacks Against Deep Learning-based Static Malware Detectors with Adversarial Byte-Level Language Model
- Learning Sequence Encoders for Temporal Knowledge Graph Completion
- Character-Word LSTM Language Models
- Word Mover's Embedding: From Word2Vec to Document Embedding
- Neural Morphological Tagging from Characters for Morphologically Rich Languages
- Learning in Text Streams: Discovery and Disambiguation of Entity and Relation Instances
- Review Helpfulness Assessment based on Convolutional Neural Network
- Implicitly Incorporating Morphological Information into Word Embedding
- MICE: Mining Idioms with Contextual Embeddings
- Ruminating Reader: Reasoning with Gated Multi-Hop Attention
- Bridging the Gap for Tokenizer-Free Language Models
- Enriching Rare Word Representations in Neural Language Models by Embedding Matrix Augmentation
- PharmKE: Knowledge Extraction Platform for Pharmaceutical Texts using Transfer Learning
- DeFINE: DEep Factorized INput Token Embeddings for Neural Sequence Modeling
- WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets
- End-to-end deep meta modelling to calibrate and optimize energy consumption and comfort
- A Neural Language Model for Dynamically Representing the Meanings of Unknown Words and Entities in a Discourse
- Characterizing machine learning process: A maturity framework
- Entity Linking Meets Deep Learning: Techniques and Solutions
- Learning K-way D-dimensional Discrete Code For Compact Embedding Representations
- AMR Parsing via Graph-Sequence Iterative Inference
- Can Neural Networks Understand Logical Entailment?
- Radical-level Ideograph Encoder for RNN-based Sentiment Analysis of Chinese and Japanese
- AMR Parsing as Sequence-to-Graph Transduction
- Large-Scale Machine Translation between Arabic and Hebrew: Available Corpora and Initial Results
- Neural Paraphrase Identification of Questions with Noisy Pretraining
- Improving Interpretability of Word Embeddings by Generating Definition and Usage
- Unsupervised Discovery of Temporal Structure in Noisy Data with Dynamical Components Analysis
- Application of a Hybrid Bi-LSTM-CRF model to the task of Russian Named Entity Recognition
- A Financial Service Chatbot based on Deep Bidirectional Transformers
- Shallow Syntax in Deep Water
- Neural Machine Translation with Byte-Level Subwords
- Integrated Sequence Tagging for Medieval Latin Using Deep Representation Learning
- Language Detection Engine for Multilingual Texting on Mobile Devices
- Unsupervised Neural Hidden Markov Models
- Generating Sentiment-Preserving Fake Online Reviews Using Neural Language Models and Their Human- and Machine-based Detection
- Compound Probabilistic Context-Free Grammars for Grammar Induction
- Review Helpfulness Prediction with Embedding-Gated CNN
- Word Shape Matters: Robust Machine Translation with Visual Embedding
- Character-Level Language Modeling with Hierarchical Recurrent Neural Networks
- Unsupervised and Efficient Vocabulary Expansion for Recurrent Neural Network Language Models in ASR
- Few-Shot Representation Learning for Out-Of-Vocabulary Words
- Dance Dance Convolution
- Personalized neural language models for real-world query auto completion
- Broad-Coverage Semantic Parsing as Transduction
- BPE and CharCNNs for Translation of Morphology: A Cross-Lingual Comparison and Analysis
- Comparative Evaluation of Pretrained Transfer Learning Models on Automatic Short Answer Grading
- Effects of Loss Functions And Target Representations on Adversarial Robustness
- A General-Purpose Tagger with Convolutional Neural Networks
- Evolving Character-level Convolutional Neural Networks for Text Classification
- Counterfactual Language Model Adaptation for Suggesting Phrases
- A Neural Network Approach for Mixing Language Models
- Discovery of Natural Language Concepts in Individual Units of CNNs
- Exploiting Synchronized Lyrics And Vocal Features For Music Emotion Detection
- Teaching Machines to Converse
- Learning from Fact-checkers: Analysis and Generation of Fact-checking Language
- A Survey on Understanding, Visualizations, and Explanation of Deep Neural Networks
- A Survey On Neural Word Embeddings
- Batch-normalized Recurrent Highway Networks
- End-to-end deep metamodeling to calibrate and optimize energy loads
- On the Importance of Word and Sentence Representation Learning in Implicit Discourse Relation Classification
- End-to-End Text Classification via Image-based Embedding using Character-level Networks
- Trends and Advancements in Deep Neural Network Communication
- Second-Order Word Embeddings from Nearest Neighbor Topological Features
- Neural Multi-Step Reasoning for Question Answering on Semi-Structured Tables
- Generic Multilayer Network Data Analysis with the Fusion of Content and Structure
- How Large a Vocabulary Does Text Classification Need? A Variational Approach to Vocabulary Selection
- A Convolutional Neural Network for Language-Agnostic Source Code Summarization
- Taking Notes on the Fly Helps BERT Pre-training
- Embedding Symbolic Knowledge into Deep Networks
- Restoring ancient text using deep learning: a case study on Greek epigraphy
- Improving Pre-Trained Multilingual Models with Vocabulary Expansion
- The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
- Automatic Arabic Dialect Identification Systems for Written Texts: A Survey
- Misspelling Correction with Pre-trained Contextual Language Model
- Learning Less-Overlapping Representations
- Effective Subword Segmentation for Text Comprehension
- Investigation of learning abilities on linguistic features in sequence-to-sequence text-to-speech synthesis
- Automated Source Code Generation and Auto-completion Using Deep Learning: Comparing and Discussing Current Language-Model-Related Approaches
- N-gram and Neural Language Models for Discriminating Similar Languages
- Predicting user intent from search queries using both CNNs and RNNs
- Character-Aware Decoder for Translation into Morphologically Rich Languages
- Syntax-Aware Language Modeling with Recurrent Neural Networks
- Bootstrapping NLU Models with Multi-task Learning
- Neural Semi-Markov Conditional Random Fields for Robust Character-Based Part-of-Speech Tagging
- Efficient Contextual Representation Learning Without Softmax Layer
- Exploring Weight Symmetry in Deep Neural Networks
- Hyperparameter optimization with REINFORCE and Transformers
- Learning Individualized Treatment Rules with Estimated Translated Inverse Propensity Score
- Neural Machine Translation: A Review and Survey
- Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NER
- Advances and Challenges in Deep Lip Reading
- Understanding Chat Messages for Sticker Recommendation in Messaging Apps
- A Multifaceted Evaluation of Neural versus Phrase-Based Machine Translation for 9 Language Directions
- A character representation enhanced on-device Intent Classification
- Long Short-Term Memory with Dynamic Skip Connections
- Evaluating Neural Word Embeddings for Sanskrit
- Multi-view Subword Regularization
- Shift-Reduce Constituent Parsing with Neural Lookahead Features
- Indicatements that character language models learn English morpho-syntactic units and regularities
- SoPa: Bridging CNNs, RNNs, and Weighted Finite-State Machines
- Character n-gram Embeddings to Improve RNN Language Models
- Auto-ML Deep Learning for Rashi Scripts OCR
- Embedding Symbolic Temporal Knowledge into Deep Sequential Models
- A Transition-based Parser for Unscoped Episodic Logical Forms
- Enhancing Sentence Relation Modeling with Auxiliary Character-level Embedding
- Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance Approach
- Leveraging Sequence Embedding and Convolutional Neural Network for Protein Function Prediction
- What do character-level models learn about morphology? The case of dependency parsing
- An Investigation of the Interactions Between Pre-Trained Word Embeddings, Character Models and POS Tags in Dependency Parsing
- Mapping high-performance RNNs to in-memory neuromorphic chips
- Revisiting Robust Neural Machine Translation: A Transformer Case Study
- Cross-Lingual Transfer for Distantly Supervised and Low-resources Indonesian NER
- UMDSub at SemEval-2018 Task 2: Multilingual Emoji Prediction Multi-channel Convolutional Neural Network on Subword Embedding
- Enhancing Chinese Intent Classification by Dynamically Integrating Character Features into Word Embeddings with Ensemble Techniques
- Improving Character-level Japanese-Chinese Neural Machine Translation with Radicals as an Additional Input Feature
- Multi-lingual Common Semantic Space Construction via Cluster-consistent Word Embedding
- What's Going On in Neural Constituency Parsers? An Analysis
- A Deep Structural Model for Analyzing Correlated Multivariate Time Series
- Byte-Level Recursive Convolutional Auto-Encoder for Text
- Assessing the Memory Ability of Recurrent Neural Networks
- Adaptive Embedding Gate for Attention-Based Scene Text Recognition
- Pretrained language model transfer on neural named entity recognition in Indonesian conversational texts
- A Multilingual Encoding Method for Text Classification and Dialect Identification Using Convolutional Neural Network
- Graph Convolution for Multimodal Information Extraction from Visually Rich Documents
- Benchmarking Approximate Inference Methods for Neural Structured Prediction
- Knowledge Distillation For Recurrent Neural Network Language Modeling With Trust Regularization
- Gating Mechanisms for Combining Character and Word-level Word Representations: An Empirical Study
- Joint Semantic Synthesis and Morphological Analysis of the Derived Word
- Prototypical Recurrent Unit
- Multimodal Embeddings from Language Models
- Nonsymbolic Text Representation
- Articulation rate in Swedish child-directed speech increases as a function of the age of the child even when surprisal is controlled for
- Transferable Natural Language Interface to Structured Queries aided by Adversarial Generation
- Deep Multimodal Learning: An Effective Method for Video Classification
- Character-Level Models versus Morphology in Semantic Role Labeling
- Quantifying Novelty and Influence, and the Patterns of Paradigm Shifts
- Char2Subword: Extending the Subword Embedding Space Using Robust Character Compositionality
- Deep N-ary Error Correcting Output Codes
- Thick-Net: Parallel Network Structure for Sequential Modeling
- Misspelling Oblivious Word Embeddings
- Improving Neural Language Models by Segmenting, Attending, and Predicting the Future
- Subword Language Model for Query Auto-Completion
- All Word Embeddings from One Embedding
- CRNN: A Joint Neural Network for Redundancy Detection
- Character-Aware Attention-Based End-to-End Speech Recognition
- Parsimonious Morpheme Segmentation with an Application to Enriching Word Embeddings
- Tabula nearly rasa: Probing the Linguistic Knowledge of Character-Level Neural Language Models Trained on Unsegmented Text
- Sensei: Self-Supervised Sensor Name Segmentation
- Enriching Under-Represented Named-Entities To Improve Speech Recognition Performance
- Learning a Word-Level Language Model with Sentence-Level Noise Contrastive Estimation for Contextual Sentence Probability Estimation
- Sent2Matrix: Folding Character Sequences in Serpentine Manifolds for Two-Dimensional Sentence
- Learning Representations for Zero-Shot Retrieval over Structured Data
- Allocating Large Vocabulary Capacity for Cross-lingual Language Model Pre-training
- Vulgaris: Analysis of a Corpus for Middle-Age Varieties of Italian Language
- Method and Dataset Entity Mining in Scientific Literature: A CNN + Bi-LSTM Model with Self-attention
- Integrating Approaches to Word Representation
- Linguistic Typology Features from Text: Inferring the Sparse Features of World Atlas of Language Structures
- Patterns versus Characters in Subword-aware Neural Language Modeling
- Compositional Sentence Representation from Character within Large Context Text
- Feature Learning for Meta-Paths in Knowledge Graphs
- DeepVar: An End-to-End Deep Learning Approach for Genomic Variant Recognition in Biomedical Literature
- Obfuscation for Privacy-preserving Syntactic Parsing
- Language modeling with Neural trans-dimensional random fields
- Learning neural trans-dimensional random field language models with noise-contrastive estimation
- A Survey on Green Deep Learning
- A Graph Neural Network Approach for Scalable and Dynamic IP Similarity in Enterprise Networks
- Improving Target-side Lexical Transfer in Multilingual Neural Machine Translation
- Paradigm Shift in Language Modeling: Revisiting CNN for Modeling Sanskrit Originated Bengali and Hindi Language
- ODSQA: Open-domain Spoken Question Answering Dataset
- Effective Character-augmented Word Embedding for Machine Reading Comprehension
- Combining Deep Learning and String Kernels for the Localization of Swiss German Tweets
- R-grams: Unsupervised Learning of Semantic Units in Natural Language
- A Computational Theory for Life-Long Learning of Semantics
- Robust Open-Vocabulary Translation from Visual Text Representations
- Phonetic-and-Semantic Embedding of Spoken Words with Applications in Spoken Content Retrieval
- Character-level Representations Improve DRS-based Semantic Parsing Even in the Age of BERT
- Stem-driven Language Models for Morphologically Rich Languages
- Revisiting Neural Language Modelling with Syllables
- The Foundations of Deep Learning with a Path Towards General Intelligence