Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages
Hannah Liu, Junghyun Min, En-Shiun Annie Lee +12
Despite major advances in machine translation (MT) in recent years, progress remains limited for many low-resource languages that lack large-scale training data and linguistic reso…
cs.CL2026
OasisSimp: An Open-source Asian-English Sentence Simplification Dataset
Hannah Liu, Muxin Tian, Iqra Ali +8
Sentence simplification aims to make complex text more accessible by reducing linguistic complexity while preserving the original meaning. However, progress in this area remains li…
cs.CL2024
SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects
David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen +5
Despite the progress we have recorded in the last few years in multilingual natural language processing, evaluation is typically limited to a small set of languages with available…