collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

Will Hawkins, Kaivalya Rawal, Jonathan Rystrøm +8

Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown that this increase in capability…

cs.CL2026

Grounding Text Embeddings in Stakeholder Associations

Jonathan Rystrøm, Sofie Burgos-Thorsen, Zihao Fu +3

Text embeddings are widely used to analyse large corpora of complex texts. However, it is unclear whether the embeddings capture the same semantic distances as the human experts us…

cs.CL2025

MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83

Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…

cs.CL2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…

cs.CL2025

Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs

Jonathan Rystrøm, Hannah Rose Kirk, Scott Hale

Large Language Models (LLMs) are becoming increasingly capable across global languages. However, the ability to communicate across languages does not necessarily translate to appro…