9 papers
MVEB: Massive Video Embedding Benchmark
Adnan El Assadi, Roman Solomatin, Isaac Chung +13
We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classificati…
Naturalistic measure of social norms alignment
Yevhen Kostiuk, Kenneth Enevoldsen, Peter Bjerregaard Vahlstrup +2
Social norms reflect shared expectations on acceptable behavior. Measuring social norms alignment remains challenging, with existing approaches typically relying on artificial clos…
Improving reasoning at inference time via uncertainty minimisation
Nicolas Legrand, Kenneth Enevoldsen, Márton Kardos +1
Large language models (LLMs) now exhibit strong multi-step reasoning abilities, but existing inference-time scaling methods remain computationally expensive, often relying on exten…
MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83
Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…
Dynaword: From One-shot to Continuously Developed Datasets
Kenneth Enevoldsen, Kristian Nørgaard Jensen, Jan Kostkan +14
Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguousl…
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
Isaac Chung, Imene Kerboua, Marton Kardos +2
The Massive Text Embedding Benchmark (MTEB) has become a standard evaluation platform for text embedding models. While previous work has established the core benchmark methodology,…