activity
20212025
collaborators

6 papers

cs.CL2025

The PLLuM Instruction Corpus

Piotr Pęzik, Filip Żarnecki, Konrad Kaczyński +50

This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project.…

cs.CL2025

PLLuM: A Family of Polish Large Language Models

Jan Kocoń, Maciej Piasecki, Arkadiusz Janz +96

Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for ot…

cs.CL2025

Integrating gender inclusivity into large language models via instruction tuning

Alina Wróblewska, Bartosz Żuk

Imagine a language with masculine, feminine, and neuter grammatical genders, yet, due to historical and political conventions, masculine forms are predominantly used to refer to me…

cs.CL2024

Investigating large language models for their competence in extracting grammatically sound sentences from transcribed noisy utterances

Alina Wróblewska

Selectively processing noisy utterances while effectively disregarding speech-specific elements poses no considerable challenge for humans, as they exhibit remarkable cognitive abi…

cs.CL2021

COMBO: State-of-the-Art Morphosyntactic Analysis

Mateusz Klimaszewski, Alina Wróblewska

We introduce COMBO - a fully neural NLP system for accurate part-of-speech tagging, morphological analysis, lemmatisation, and (enhanced) dependency parsing. It predicts categorica…

cs.CL2021

COMBO: a new module for EUD parsing

Mateusz Klimaszewski, Alina Wróblewska

We introduce the COMBO-based approach for EUD parsing and its implementation, which took part in the IWPT 2021 EUD shared task. The goal of this task is to parse raw texts in 17 la…