Low-Resource Named Entity Recognition with Cross-Lingual, Character-Level Neural Conditional Random Fields
arXiv:2404.09383
Abstract
Low-resource named entity recognition is still an open problem in NLP. Most state-of-the-art systems require tens of thousands of annotated sentences in order to obtain high performance. However, for most of the world's languages, it is unfeasible to obtain such annotation. In this paper, we present a transfer learning scheme, whereby we train character-level neural CRFs to predict named entities for both high-resource languages and low resource languages jointly. Learning character representations for multiple related languages allows transfer among the languages, improving F1 by up to 9.8 points over a loglinear CRF baseline.
IJCNLP 2017
References in corpus (4)
Cited by in corpus (22)
- Few-shot classification in Named Entity Recognition Task
- Neural Entity Linking: A Survey of Models Based on Deep Learning
- Low-resource Languages: A Review of Past Work and Future Challenges
- Evaluating Language Model Finetuning Techniques for Low-resource Languages
- Localization of Fake News Detection via Multitask Transfer Learning
- Neural Cross-Lingual Named Entity Recognition with Minimal Resources
- On Difficulties of Cross-Lingual Transfer with Order Differences: A Case Study on Dependency Parsing
- Chinese Discourse Segmentation Using Bilingual Discourse Commonality
- Dependency-Guided LSTM-CRF for Named Entity Recognition
- Augmented Natural Language for Generative Sequence Labeling
- KINNEWS and KIRNEWS: Benchmarking Cross-Lingual Text Classification for Kinyarwanda and Kirundi
- Global Attention for Name Tagging
- Low-Resource Sequence Labeling via Unsupervised Multilingual Contextualized Representations
- What Matters for Neural Cross-Lingual Named Entity Recognition: An Empirical Analysis
- Cross-Lingual Transfer for Distantly Supervised and Low-resources Indonesian NER
- Exploiting Cross-Lingual Subword Similarities in Low-Resource Document Classification
- Open Named Entity Modeling from Embedding Distribution
- Multi-task Learning Based Neural Bridging Reference Resolution
- XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word Alignment
- Free the Plural: Unrestricted Split-Antecedent Anaphora Resolution
- Do "English" Named Entity Recognizers Work Well on Global Englishes?
- Intent Classification and Slot Filling for Privacy Policies