3 papers
cs.CL2025
Employing Sentence Space Embedding for Classification of Data Stream from Fake News Domain
PaweÅ Zyblewski, Jakub Klikowski, Weronika Borek-Marciniec +1
Tabular data is considered the last unconquered castle of deep learning, yet the task of data stream classification is stated to be an equally important and demanding research area…
cs.CL2025
Large Language Models in Legislative Content Analysis: A Dataset from the Polish Parliament
Arkadiusz BryÅkowski, Jakub Klikowski
Large language models (LLMs) are among the best methods for processing natural language, partly due to their versatility. At the same time, domain-specific LLMs are more practical…
cs.CL2024
WarCov -- Large multilabel and multimodal dataset from social platform
Weronika Borek-Marciniec, Pawel Zyblewski, Jakub Klikowski +1
In the classification tasks, from raw data acquisition to the curation of a dataset suitable for use in evaluating machine learning models, a series of steps - often associated wit…