TAPE: Assessing Few-shot Russian Language Understanding
arXiv:2210.12813 · doi:10.18653/v1/2022.findings-emnlp.183
Abstract
Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes. However, this fast-growing area lacks standardized evaluation suites for non-English languages, hindering progress outside the Anglo-centric paradigm. To address this line of research, we propose TAPE (Text Attack and Perturbation Evaluation), a novel benchmark that includes six more complex NLU tasks for Russian, covering multi-hop reasoning, ethical concepts, logic and commonsense knowledge. The TAPE's design focuses on systematic zero-shot and few-shot NLU evaluation: (i) linguistic-oriented adversarial attacks and perturbations for analyzing robustness, and (ii) subpopulations for nuanced interpretation. The detailed analysis of testing the autoregressive baselines indicates that simple spelling-based perturbations affect the performance the most, while paraphrasing the input has a more negligible effect. At the same time, the results demonstrate a significant gap between the neural and human baselines for most tasks. We publicly release TAPE (tape-benchmark.com) to foster research on robust LMs that can generalize to new tasks when little to no supervision is available.
Accepted to EMNLP 2022 Findings
References in corpus (14)
- Scikit-learn: Machine Learning in Python
- BERTScore: Evaluating Text Generation with BERT
- SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- True Few-Shot Learning with Language Models
- Aligning AI With Shared Human Values
- Few-shot Learning with Multilingual Language Models
- Mapping global dynamics of benchmark creation and saturation in artificial intelligence
- Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
- WorldTree: A Corpus of Explanation Graphs for Elementary Science Questions supporting Multi-Hop Inference
- FewCLUE: A Chinese Few-shot Learning Evaluation Benchmark
- Data Augmentation for Low-Resource Named Entity Recognition Using Backtranslation
- FewNLU: Benchmarking State-of-the-Art Methods for Few-Shot Natural Language Understanding