4 citations
6 papers · 1 filter
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
Qiwei Peng, Yekun Chai, Xuhong Li
Large language models (LLMs) have made significant progress in generating codes from textual prompts. However, existing benchmarks have mainly concentrated on translating English p…
A Multilingual Perspective on Probing Gender Bias
Karolina Stańczak
Gender bias represents a form of systematic negative treatment that targets individuals based on their gender. This discrimination can range from subtle sexist remarks and gendered…
Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language Models
Phillip Rust, Anders Søgaard
Language models such as mBERT, XLM-R, and BLOOM aim to achieve multilingual generalization or compression to facilitate transfer to a large number of (potentially unseen) languages…
Are Pretrained Multilingual Models Equally Fair Across Languages?
Laura Cabello Piqueras, Anders Søgaard
Pretrained multilingual language models can help bridge the digital language divide, enabling high-quality NLP models for lower resourced languages. Studies of multilingual models…
Machine Reading, Fast and Slow: When Do Models "Understand" Language?
Sagnik Ray Choudhury, Anna Rogers, Isabelle Augenstein
Two of the most fundamental challenges in Natural Language Understanding (NLU) at present are: (a) how to establish whether deep learning-based models score highly on NLU benchmark…
On Training Instance Selection for Few-Shot Neural Text Generation
Ernie Chang, Xiaoyu Shen, Hui-Syuan Yeh +1
Large-scale pretrained language models have led to dramatic improvements in text generation. Impressive performance can be achieved by finetuning only on a small number of instance…