7 citations · 7 across the 3 of their papers we have counts for
8 papers
The PLLuM Instruction Corpus
Piotr Pęzik, Filip Żarnecki, Konrad Kaczyński +50
This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project.…
PLLuM: A Family of Polish Large Language Models
Jan Kocoń, Maciej Piasecki, Arkadiusz Janz +96
Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for ot…
SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata
This article introduces semantically meaningful causal language modeling (SMCLM), a selfsupervised method of training autoregressive models to generate semantically equivalent text…
Unveiling Dual Quality in Product Reviews: An NLP-Based Approach
Rafał Poświata, Marcin Michał Mirończuk, Sławomir Dadas +2
Consumers often face inconsistent product quality, particularly when identical products vary between markets, a situation known as the dual quality problem. To identify and address…
Evaluating Polish linguistic and cultural competency in large language models
Sławomir Dadas, Małgorzata Grębowiec, Michał Perełkiewicz +1
Large language models (LLMs) are becoming increasingly proficient in processing and generating multilingual texts, which allows them to address real-world problems more effectively…
MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83
Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…