3 papers
cs.CL2025
MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83
Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…
cs.CL2025
Dynaword: From One-shot to Continuously Developed Datasets
Kenneth Enevoldsen, Kristian Nørgaard Jensen, Jan Kostkan +14
Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguousl…
cs.LG2025
-- Semantic Signal Separation
Márton Kardos, Jan Kostkan, Arnault-Quentin Vermillet +3
Topic models are useful tools for discovering latent semantic structures in large textual corpora. Recent efforts have been oriented at incorporating contextual representations in…