activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

One Script Instead of Hundreds? On Pretraining Romanized Encoder Language Models

Benedikt Ebing, Lennart Keller, Goran Glavaš

Exposing latent lexical overlap, script romanization has emerged as an effective strategy for improving cross-lingual transfer (XLT) in multilingual language models (mLMs). Most pr…

cs.CL2025

TransAlign: Machine Translation Encoders are Strong Word Aligners, Too

Benedikt Ebing, Christian Goldschmied, Goran Glavaš

In the absence of sizable training data for most world languages and NLP tasks, translation-based strategies such as translate-test -- evaluating on noisy source language data tran…

cs.CL2025

The Devil Is in the Word Alignment Details: On Translation-Based Cross-Lingual Transfer for Token Classification Tasks

Benedikt Ebing, Goran Glavaš

Translation-based strategies for cross-lingual transfer XLT such as translate-train -- training on noisy target language data translated from the source language -- and translate-t…

cs.CL2025

ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding

Indraneil Paul, Haoyi Yang, Goran Glavaš +2

Language models (LMs) have become a staple of the code-writing toolbox. Their pre-training recipe has, however, remained stagnant over recent years, barring the occasional changes…

cs.CL2024

SpeechTaxi: On Multilingual Semantic Speech Classification

Lennart Keller, Goran Glavaš

Recent advancements in multilingual speech encoding as well as transcription raise the question of the most effective approach to semantic speech classification. Concretely, can (1…

cs.CL2024

To Translate or Not to Translate: A Systematic Investigation of Translation-Based Cross-Lingual Transfer to Low-Resource Languages

Benedikt Ebing, Goran Glavaš

Perfect machine translation (MT) would render cross-lingual transfer (XLT) by means of multilingual language models (mLMs) superfluous. Given, on the one hand, the large body of wo…