9 papers
Improving reasoning at inference time via uncertainty minimisation
Nicolas Legrand, Kenneth Enevoldsen, Márton Kardos +1
Large language models (LLMs) now exhibit strong multi-step reasoning abilities, but existing inference-time scaling methods remain computationally expensive, often relying on exten…
MAEB: Massive Audio Embedding Benchmark
Adnan El Assadi, Isaac Chung, Chenghao Xiao +15
We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasonin…
HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
Adnan El Assadi, Isaac Chung, Roman Solomatin +2
Comparing human and model performance offers a valuable perspective for understanding the strengths and limitations of embedding models, highlighting where they succeed and where t…
Dynaword: From One-shot to Continuously Developed Datasets
Kenneth Enevoldsen, Kristian Nørgaard Jensen, Jan Kostkan +14
Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguousl…
Continuous sentiment scores for literary and multilingual contexts
Laurits Lyngbaek, Pascale Feldkamp, Yuri Bizzoni +2
Sentiment Analysis is widely used to quantify sentiment in text, but its application to literary texts poses unique challenges due to figurative language, stylistic ambiguity, as w…
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
Isaac Chung, Imene Kerboua, Marton Kardos +2
The Massive Text Embedding Benchmark (MTEB) has become a standard evaluation platform for text embedding models. While previous work has established the core benchmark methodology,…