BPEmb: Tokenization-free Pre-trained Subword Embeddings in 275 Languages
arXiv:1710.02187
Abstract
We present BPEmb, a collection of pre-trained subword unit embeddings in 275 languages, based on Byte-Pair Encoding (BPE). In an evaluation using fine-grained entity typing as testbed, BPEmb performs competitively, and for some languages bet- ter than alternative subword approaches, while requiring vastly fewer resources and no tokenization. BPEmb is available at https://github.com/bheinzerling/bpemb
Cited by in corpus (4)
- NLNDE: Enhancing Neural Sequence Taggers with Attention and Noisy Channel for Robust Pharmacological Entity Detection
- On the Impact of Knowledge-based Linguistic Annotations in the Quality of Scientific Embeddings
- Unsupervised Domain Adaptation of a Pretrained Cross-Lingual Language Model
- Vec2Sent: Probing Sentence Embeddings with Natural Language Generation