2 papers
cs.CL2025
MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83
Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…
cs.CL2025
The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design
Artem Snegirev, Maria Tikhonova, Anna Maksimova +2
Embedding models play a crucial role in Natural Language Processing (NLP) by creating text embeddings used in various tasks such as information retrieval and assessing semantic tex…