8 papers
A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books
Varun Ghat Ravikumar, Sina Ahmadi, Lena Jäger +1
Most endangered languages lack the parallel data required for machine translation, despite the existence of descriptive grammar books. We introduce a pipeline that uses large langu…
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
Negar Foroutan, Clara Meister, Debjit Paul +4
Tokenization is the first -- and often least scrutinized -- step of most NLP pipelines. Standard algorithms for learning tokenizers rely on frequency-based objectives, which favor…
ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs
Andrianos Michail, Stylianos Psychias, Michelle Wastl +3
Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a limited set of languages, ar…
Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias
Elias Schuhmacher, Andrianos Michail, Juri Opitz +2
To be discoverable in an embedding-based search process, each part of a document should be reflected in its embedding representation. To quantify any potential reflection biases, w…
CommonMorph: Participatory Morphological Documentation Platform
Aso Mahmudi, Sina Ahmadi, Kemal Kurniawan +3
Collecting and annotating morphological data present significant challenges, requiring linguistic expertise, methodological rigour, and substantial resources. These barriers are pa…
Meaningful Pose-Based Sign Language Evaluation
Zifan Jiang, Colin Leong, Amit Moryossef +8
We present a comprehensive study on meaningfully evaluating sign language utterances in the form of human skeletal poses. The study covers keypoint distance-based, embedding-based,…