collaborators

5 papers

cs.CL2025

Dynaword: From One-shot to Continuously Developed Datasets

Kenneth Enevoldsen, Kristian Nørgaard Jensen, Jan Kostkan +14

Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguousl…

cs.CL2025

Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks

Isaac Chung, Imene Kerboua, Marton Kardos +2

The Massive Text Embedding Benchmark (MTEB) has become a standard evaluation platform for text embedding models. While previous work has established the core benchmark methodology,…

cs.CL2025

topicwizard -- a Modern, Model-agnostic Framework for Topic Model Visualization and Interpretation

Márton Kardos, Kenneth C. Enevoldsen, Kristoffer Laigaard Nielbo

Topic models are statistical tools that allow their users to gain qualitative and quantitative insights into the contents of textual corpora without the need for close reading. The…

cs.CV2025

MIEB: Massive Image Embedding Benchmark

Chenghao Xiao, Isaac Chung, Imene Kerboua +7

Image representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whet…

cs.CL2025

MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83

Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…