CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation
arXiv:2103.06874 · doi:10.1162/tacl_a_00448
Abstract
Pipelined NLP systems have largely been superseded by end-to-end neural modeling, yet nearly all commonly-used models still require an explicit tokenization step. While recent tokenization approaches based on data-derived subword lexicons are less brittle than manually engineered tokenizers, these techniques are not equally suited to all languages, and the use of any fixed vocabulary may limit a model's ability to adapt. In this paper, we present CANINE, a neural encoder that operates directly on character sequences, without explicit tokenization or vocabulary, and a pre-training strategy that operates either directly on characters or optionally uses subwords as a soft inductive bias. To use its finer-grained input effectively and efficiently, CANINE combines downsampling, which reduces the input sequence length, with a deep transformer stack, which encodes context. CANINE outperforms a comparable mBERT model by 2.8 F1 on TyDi QA, a challenging multilingual benchmark, despite having 28% fewer model parameters.
TACL Final Version
References in corpus (12)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Cross-lingual Language Model Pretraining
- Introduction to the CoNLL-2002 Shared Task: Language-Independent Named Entity Recognition
- Hierarchical Multiscale Recurrent Neural Networks
- Random Feature Attention
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing
- Which Encoding is the Best for Text Classification in Chinese, English, Japanese and Korean?
- Hash Embeddings for Efficient Word Representations
- CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters
- English Intermediate-Task Training Improves Zero-Shot Cross-Lingual Transfer Too
- Multiscale sequence modeling with a learned dictionary
- Multi-view Subword Regularization
Cited by in corpus (17)
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
- Enhancing Phenotype Recognition in Clinical Notes Using Large Language Models: PhenoBCBERT and PhenoGPT
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Language Modelling with Pixels
- Specializing Multilingual Language Models: An Empirical Study
- Efficient Transformers with Dynamic Token Pooling
- A Survey of Text Representation Methods and Their Genealogy
- ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
- A multimodal deep learning architecture for smoking detection with a small data approach
- Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation
- An Information Extraction Study: Take In Mind the Tokenization!
- Learning to Look Inside: Augmenting Token-Based Encoders with Character-Level Information
- Predictability and Causality in Spanish and English Natural Language Generation
- You should evaluate your language model on marginal likelihood over tokenisations
- Optimal word order for non-causal text generation with Large Language Models: the Spanish case
- Integrating Approaches to Word Representation