5 papers
Adapting LLMs for Minimal-edit Grammatical Error Correction
Ryszard Staruch, Filip Graliński, Daniel Dzienisiewicz
Decoder-only large language models have shown superior performance in the fluency-edit English Grammatical Error Correction, but their adaptation for minimal-edit English GEC is st…
LLMzSzŁ: a comprehensive LLM benchmark for Polish
Krzysztof Jassem, Michał Ciesiółka, Filip Graliński +5
This article introduces the first comprehensive benchmark for the Polish language at this scale: LLMzSzŁ (LLMs Behind the School Desk). It is based on a coherent collection of Poli…
Oddballness: universal anomaly detection with language models
Filip Graliński, Ryszard Staruch, Krzysztof Jurkiewicz
We present a new method to detect anomalies in texts (in general: in sequences of any data), using language models, in a totally unsupervised manner. The method considers probabili…
POLygraph: Polish Fake News Dataset
Daniel Dzienisiewicz, Filip Graliński, Piotr Jabłoński +3
This paper presents the POLygraph dataset, a unique resource for fake news detection in Polish. The dataset, created by an interdisciplinary team, is composed of two parts: the "fa…
Two Approaches to Diachronic Normalization of Polish Texts
Kacper Dudzic, Filip Graliński, Krzysztof Jassem +2
This paper discusses two approaches to the diachronic normalization of Polish texts: a rule-based solution that relies on a set of handcrafted patterns, and a neural normalization…