3 papers
cs.CL2024
POLygraph: Polish Fake News Dataset
Daniel Dzienisiewicz, Filip Graliński, Piotr Jabłoński +3
This paper presents the POLygraph dataset, a unique resource for fake news detection in Polish. The dataset, created by an interdisciplinary team, is composed of two parts: the "fa…
cs.CL2024
Two Approaches to Diachronic Normalization of Polish Texts
Kacper Dudzic, Filip Graliński, Krzysztof Jassem +2
This paper discusses two approaches to the diachronic normalization of Polish texts: a rule-based solution that relies on a set of handcrafted patterns, and a neural normalization…
cs.CL2023
Back Transcription as a Method for Evaluating Robustness of Natural Language Understanding Models to Speech Recognition Errors
Marek Kubis, Paweł Skórzewski, Marcin Sowański +1
In a spoken dialogue system, an NLU model is preceded by a speech recognition system that can deteriorate the performance of natural language understanding. This paper proposes a m…