4 papers · 1 filter
IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages
Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare +14
Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and lang…
DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information
Roland Roller, Vera Czehmann, Derya Erman +13
Conversational data collected in domains such as healthcare or social sciences is a valuable resource for research and automated analysis. However, responsible data sharing require…
Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures
Chiara Manna, Hosein Mohebbi, Afra Alishahi +2
While Large Language Models achieve state-of-the-art results across a wide range of NLP tasks, they remain prone to systematic biases. Among these, gender bias is particularly sali…
What do Large Language Models Need for Machine Translation Evaluation?
Shenbin Qian, Archchana Sindhujan, Minnie Kabra +4
Leveraging large language models (LLMs) for various natural language processing tasks has led to superlative claims about their performance. For the evaluation of machine translati…