6 papers
PL-MTEB: Polish Massive Text Embedding Benchmark
RafaÅ PoÅwiata, SÅawomir Dadas, MichaÅ PereÅkiewicz
In this paper, we introduce the Polish Massive Text Embedding Benchmark (PL-MTEB), a comprehensive benchmark for text embeddings in the Polish language. PL-MTEB comprises 30 divers…
Long-Context Encoder Models for Polish Language Understanding
SÅawomir Dadas, RafaÅ PoÅwiata, Marek KozÅowski +4
While decoder-only Large Language Models (LLMs) have recently dominated the NLP landscape, encoder-only architectures remain a cost-effective and parameter-efficient standard for d…
PLLuM: A Family of Polish Large Language Models
Jan KocoÅ, Maciej Piasecki, Arkadiusz Janz +96
Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for ot…
SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
MichaÅ PereÅkiewicz, SÅawomir Dadas, RafaÅ PoÅwiata
This article introduces semantically meaningful causal language modeling (SMCLM), a selfsupervised method of training autoregressive models to generate semantically equivalent text…
Unveiling Dual Quality in Product Reviews: An NLP-Based Approach
RafaÅ PoÅwiata, Marcin MichaÅ MiroÅczuk, SÅawomir Dadas +2
Consumers often face inconsistent product quality, particularly when identical products vary between markets, a situation known as the dual quality problem. To identify and address…
Evaluating Polish linguistic and cultural competency in large language models
SÅawomir Dadas, MaÅgorzata GrÄbowiec, MichaÅ PereÅkiewicz +1
Large language models (LLMs) are becoming increasingly proficient in processing and generating multilingual texts, which allows them to address real-world problems more effectively…