collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

Sławomir Dadas, Michał Perełkiewicz, Rafał Poświata +3

Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understanding and generation have bee…

cs.CL2026

PL-MTEB: Polish Massive Text Embedding Benchmark

Rafał Poświata, Sławomir Dadas, Michał Perełkiewicz

In this paper, we introduce the Polish Massive Text Embedding Benchmark (PL-MTEB), a comprehensive benchmark for text embeddings in the Polish language. PL-MTEB comprises 30 divers…

cs.CL2026

Long-Context Encoder Models for Polish Language Understanding

Sławomir Dadas, Rafał Poświata, Marek Kozłowski +4

While decoder-only Large Language Models (LLMs) have recently dominated the NLP landscape, encoder-only architectures remain a cost-effective and parameter-efficient standard for d…

cs.CL2025

PLLuM: A Family of Polish Large Language Models

Jan Kocoń, Maciej Piasecki, Arkadiusz Janz +96

Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for ot…

cs.CL2025

SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation

Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata

This article introduces semantically meaningful causal language modeling (SMCLM), a selfsupervised method of training autoregressive models to generate semantically equivalent text…

cs.CL2025

Unveiling Dual Quality in Product Reviews: An NLP-Based Approach

Rafał Poświata, Marcin Michał Mirończuk, Sławomir Dadas +2

Consumers often face inconsistent product quality, particularly when identical products vary between markets, a situation known as the dual quality problem. To identify and address…