activity
20202026
most citedSubword Segmental Language Modelling for Nguni Languages

1 citations · 2 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

The Geometry of Low-Resource Language Representations

Francois Meyer, Jan Buys

The performance gap between low- and high-resource languages in LLMs is widely known, but it remains unclear which internal model factors drive these disparities. In this paper, we…

cs.CL2026

MzansiText and MzansiLM: An Open Corpus and Decoder-Only Language Model for South African Languages

Anri Lombard, Simbarashe Mawere, Temi Aina +5

Decoder-only language models can be adapted to diverse tasks through instruction finetuning, but the extent to which this generalizes at small scale for low-resource languages rema…

cs.CL2025

The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages

Francois Meyer, Jan Buys

Subword segmentation is typically applied in preprocessing and stays fixed during training. Alternatively, it can be learned during training to optimise the training objective. In…

cs.CL2024

A Systematic Analysis of Subwords and Cross-Lingual Transfer in Multilingual Translation

Francois Meyer, Jan Buys

Multilingual modelling can improve machine translation for low-resource languages, partly through shared subword representations. This paper studies the role of subword segmentatio…

cs.CL20241 cited

Triples-to-isiXhosa (T2X): Addressing the Challenges of Low-Resource Agglutinative Data-to-Text Generation

Francois Meyer, Jan Buys

Most data-to-text datasets are for English, so the difficulties of modelling data-to-text for low-resource languages are largely unexplored. In this paper we tackle data-to-text fo…

cs.CL2023

Subword Segmental Machine Translation: Unifying Segmentation and Target Sentence Generation

Francois Meyer, Jan Buys

Subword segmenters like BPE operate as a preprocessing step in neural machine translation and other (conditional) language models. They are applied to datasets before training, so…