activity
20232026
most citedMMTEB: Massive Multilingual Text Embedding Benchmark

15 citations · 16 across the 17 of their papers we have counts for

collaborators
Showing cs.CLShow all

16 papers · 1 filter

cs.CL2026

Last Translation Benchmark

Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…

cs.CL2026

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina +2

Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced d…

cs.CL2026

Modeling semantic association in self-paced reading with language model embeddings

Sara Møller Østergaard, Kenneth Enevoldsen, Afra Alishahi +1

Semantic association between a word and its context has been identified as an important component of reading comprehension, even when word predictability is accounted for. Recent r…

cs.CL2026

Grounding Text Embeddings in Stakeholder Associations

Jonathan Rystrøm, Sofie Burgos-Thorsen, Zihao Fu +3

Text embeddings are widely used to analyse large corpora of complex texts. However, it is unclear whether the embeddings capture the same semantic distances as the human experts us…

cs.CL2026

Naturalistic measure of social norms alignment

Yevhen Kostiuk, Kenneth Enevoldsen, Peter Bjerregaard Vahlstrup +2

Social norms reflect shared expectations on acceptable behavior. Measuring social norms alignment remains challenging, with existing approaches typically relying on artificial clos…

cs.CL2026

One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation

Yevhen Kostiuk, Kenneth Enevoldsen

Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point evaluation ignores a main prob…