5 papers
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches
Teddy Ferdinan, BartÅomiej Koptyra, MikoÅaj Langner +42
While Reasoning Language Models (RLMs) are rapidly emerging as powerful tools for scientific research, their impact is primarily concentrated in "hard science" fields. The slow --…
Chunking Methods on Retrieval-Augmented Generation - Effectiveness Evaluation Against Computational Cost and Limitations
Mateusz Åmigielski, MichaÅ Rajkowski, Mateusz Zbrocki +5
Retrieval-Augmented Generation (RAG) has demonstrated significant capabilities in enhancing the performance of Large Language Models (LLMs). One of the key tasks in RAG systems is…
The PLLuM Instruction Corpus
Piotr PÄzik, Filip Å»arnecki, Konrad KaczyÅski +50
This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project.…
MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83
Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…
PLLuM: A Family of Polish Large Language Models
Jan KocoÅ, Maciej Piasecki, Arkadiusz Janz +96
Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for ot…