7 papers
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
Indraneil Paul, Falko Helm, Goran Glavaš +1
Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing…
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
Indraneil Paul, Goran GlavaÅ¡, Goran Glavaš +1
Reward models (RMs) have become an indispensable fixture of the language model (LM) post-training playbook, enabling policy alignment and test-time scaling. Research on the applica…
One Script Instead of Hundreds? On Pretraining Romanized Encoder Language Models
Benedikt Ebing, Lennart Keller, Goran Glavaš
Exposing latent lexical overlap, script romanization has emerged as an effective strategy for improving cross-lingual transfer (XLT) in multilingual language models (mLMs). Most pr…
TransAlign: Machine Translation Encoders are Strong Word Aligners, Too
Benedikt Ebing, Christian Goldschmied, Goran Glavaš
In the absence of sizable training data for most world languages and NLP tasks, translation-based strategies such as translate-test -- evaluating on noisy source language data tran…
The Devil Is in the Word Alignment Details: On Translation-Based Cross-Lingual Transfer for Token Classification Tasks
Benedikt Ebing, Goran Glavaš
Translation-based strategies for cross-lingual transfer XLT such as translate-train -- training on noisy target language data translated from the source language -- and translate-t…
Problem Solving Through Human-AI Preference-Based Cooperation
Subhabrata Dutta, Timo Kaufmann, Goran Glavaš +7
While there is a widespread belief that artificial general intelligence (AGI) -- or even superhuman AI -- is imminent, complex problems in expert domains are far from being solved.…