1 citations · 2 across the 3 of their papers we have counts for
5 papers · 1 filter
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
Isaac Chung, Imene Kerboua, Marton Kardos +2
The Massive Text Embedding Benchmark (MTEB) has become a standard evaluation platform for text embedding models. While previous work has established the core benchmark methodology,…
topicwizard -- a Modern, Model-agnostic Framework for Topic Model Visualization and Interpretation
Márton Kardos, Kenneth C. Enevoldsen, Kristoffer Laigaard Nielbo
Topic models are statistical tools that allow their users to gain qualitative and quantitative insights into the contents of textual corpora without the need for close reading. The…
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
Kenneth Enevoldsen, Márton Kardos, Niklas Muennighoff +1
The evaluation of English text embeddings has transitioned from evaluating a handful of datasets to broad coverage across many tasks through benchmarks such as MTEB. However, this…
DANSK and DaCy 2.6.0: Domain Generalization of Danish Named Entity Recognition
Kenneth Enevoldsen, Emil Trenckner Jessen, Rebekah Baglini
Named entity recognition is one of the cornerstones of Danish NLP, essential for language technology applications within both industry and research. However, Danish NER is inhibite…
Danish Foundation Models
Kenneth Enevoldsen, Lasse Hansen, Dan S. Nielsen +10
Large language models, sometimes referred to as foundation models, have transformed multiple fields of research. However, smaller languages risk falling behind due to high training…