Polyglot: Distributed Word Representations for Multilingual NLP
arXiv:1307.1662
Abstract
Distributed word representations (word embeddings) have recently contributed to competitive performance in language modeling and several NLP tasks. In this work, we train word embeddings for more than 100 languages using their corresponding Wikipedias. We quantitatively demonstrate the utility of our word embeddings by using them as the sole features for training a part of speech tagger for a subset of these languages. We find their performance to be competitive with near state-of-art methods in English, Danish and Swedish. Moreover, we investigate the semantic features captured by these embeddings through the proximity of word groupings. We will release these embeddings publicly to help researchers in the development and enhancement of multilingual applications.
10 pages, 2 figures, Proceedings of Conference on Computational Natural Language Learning CoNLL'2013
References in corpus (2)
Cited by in corpus (46)
- DeepWalk: Online Learning of Social Representations
- Massively Multilingual Word Embeddings
- Multi-Task Cross-Lingual Sequence Tagging from Scratch
- Charting the Landscape of Online Cryptocurrency Manipulation
- sense2vec - A Fast and Accurate Method for Word Sense Disambiguation In Neural Word Embeddings
- Hierarchically-Refined Label Attention Network for Sequence Labeling
- Multilingual Models for Compositional Distributed Semantics
- Separated by an Un-common Language: Towards Judgment Language Informed Vector Space Modeling
- Multilingual Hierarchical Attention Networks for Document Classification
- Cross-lingual Models of Word Embeddings: An Empirical Comparison
- Neural Probabilistic Model for Non-projective MST Parsing
- A Neural Entity Coreference Resolution Review
- POLYGLOT-NER: Massive Multilingual Named Entity Recognition
- Hierarchical Pointer Net Parsing
- Blending gradient boosted trees and neural networks for point and probabilistic forecasting of hierarchical time series
- False-Friend Detection and Entity Matching via Unsupervised Transliteration
- Robust Multilingual Part-of-Speech Tagging via Adversarial Training
- Freshman or Fresher? Quantifying the Geographic Variation of Internet Language
- Distributed Representations for Compositional Semantics
- Semi Supervised Preposition-Sense Disambiguation using Multilingual Data
- Multilingual Visual Sentiment Concept Matching
- hauWE: Hausa Words Embedding for Natural Language Processing
- Statistically Significant Detection of Linguistic Change
- Arabic Named Entity Recognition using Word Representations
- Joint Aspect and Polarity Classification for Aspect-based Sentiment Analysis with End-to-End Neural Networks
- Cross-lingual Dataless Classification for Languages with Small Wikipedia Presence
- The Best of Both Worlds: Lexical Resources To Improve Low-Resource Part-of-Speech Tagging
- Language classification from bilingual word embedding graphs
- LAWDR: Language-Agnostic Weighted Document Representations from Pre-trained Models
- Vacaspati: A Diverse Corpus of Bangla Literature
- Enhancing Word Embeddings with Knowledge Extracted from Lexical Resources
- UzBERT: pretraining a BERT model for Uzbek
- Semi-Supervised Translation with MMD Networks
- Linking Tweets with Monolingual and Cross-Lingual News using Transformed Word Embeddings
- DiaLex: A Benchmark for Evaluating Multidialectal Arabic Word Embeddings
- New word analogy corpus for exploring embeddings of Czech words
- Replicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets
- Dependency-based Hybrid Trees for Semantic Parsing
- ALL-IN-1: Short Text Classification with One Model for All Languages
- Massively Parallel Cross-Lingual Learning in Low-Resource Target Language Translation
- External Lexical Information for Multilingual Part-of-Speech Tagging
- Sparse Coding of Neural Word Embeddings for Multilingual Sequence Labeling
- Unsupervised Learning of Morphological Forests
- EEMC: Embedding Enhanced Multi-tag Classification
- When silver glitters more than gold: Bootstrapping an Italian part-of-speech tagger for Twitter
- A Factorized Model for Transitive Verbs in Compositional Distributional Semantics