6 papers
The PLLuM Instruction Corpus
Piotr Pęzik, Filip Żarnecki, Konrad Kaczyński +50
This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project.…
PLLuM: A Family of Polish Large Language Models
Jan Kocoń, Maciej Piasecki, Arkadiusz Janz +96
Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for ot…
Integrating gender inclusivity into large language models via instruction tuning
Alina Wróblewska, Bartosz Żuk
Imagine a language with masculine, feminine, and neuter grammatical genders, yet, due to historical and political conventions, masculine forms are predominantly used to refer to me…
Investigating large language models for their competence in extracting grammatically sound sentences from transcribed noisy utterances
Alina Wróblewska
Selectively processing noisy utterances while effectively disregarding speech-specific elements poses no considerable challenge for humans, as they exhibit remarkable cognitive abi…
COMBO: State-of-the-Art Morphosyntactic Analysis
Mateusz Klimaszewski, Alina Wróblewska
We introduce COMBO - a fully neural NLP system for accurate part-of-speech tagging, morphological analysis, lemmatisation, and (enhanced) dependency parsing. It predicts categorica…
COMBO: a new module for EUD parsing
Mateusz Klimaszewski, Alina Wróblewska
We introduce the COMBO-based approach for EUD parsing and its implementation, which took part in the IWPT 2021 EUD shared task. The goal of this task is to parse raw texts in 17 la…