10 papers
PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding
Sławomir Dadas, Michał Perełkiewicz, Rafał Poświata +3
Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understanding and generation have bee…
PL-MTEB: Polish Massive Text Embedding Benchmark
RafaÅ PoÅwiata, SÅawomir Dadas, MichaÅ PereÅkiewicz
In this paper, we introduce the Polish Massive Text Embedding Benchmark (PL-MTEB), a comprehensive benchmark for text embeddings in the Polish language. PL-MTEB comprises 30 divers…
Long-Context Encoder Models for Polish Language Understanding
SÅawomir Dadas, RafaÅ PoÅwiata, Marek KozÅowski +4
While decoder-only Large Language Models (LLMs) have recently dominated the NLP landscape, encoder-only architectures remain a cost-effective and parameter-efficient standard for d…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
The PLLuM Instruction Corpus
Piotr PÄzik, Filip Å»arnecki, Konrad KaczyÅski +50
This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project.…
MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83
Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…