activity
20182024
most citedQuality at a Glance: An Audit of Web-Crawled Multilingual Datasets

177 citations · 189 across the 6 of their papers we have counts for

collaborators

9 papers

cs.CL2024

German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset

Laura Mascarell, Ribin Chalumattu, Annette Rios

The advent of Large Language Models (LLMs) has led to remarkable progress on a wide range of natural language processing tasks. Despite the advances, these large-sized models still…

cs.CL2022★ 4 cited

Considerations for meaningful sign language machine translation based on glosses

Mathias Müller, Zifan Jiang, Amit Moryossef +2

Automatic sign language processing is gaining popularity in Natural Language Processing (NLP) research (Yin et al., 2021). In machine translation (MT) in particular, sign language…

cs.CL2021★ 2 cited

Evaluating the Immediate Applicability of Pose Estimation for Sign Language Recognition

Amit Moryossef, Ioannis Tsochantaridis, Joe Dinn +6

Signed languages are visual languages produced by the movement of the hands, face, and body. In this paper, we evaluate representations based on skeleton poses, as these are explai…

cs.CL2021

On Biasing Transformer Attention Towards Monotonicity

Annette Rios, Chantal Amrhein, Noëmi Aepli +1

Many sequence-to-sequence tasks in natural language processing are roughly monotonic in the alignment between source and target sequence, and previous work has facilitated or enfor…

cs.CL2021

AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource Languages

Abteen Ebrahimi, Manuel Mager, Arturo Oncevay +14

Pretrained multilingual models are able to perform cross-lingual transfer in a zero-shot setting, even for languages unseen during pretraining. However, prior work evaluating perfo…

cs.CL2021★ 177 cited

Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets

Julia Kreutzer, Isaac Caswell, Lisa Wang +49

With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, web-mined text dataset…