3 papers
cs.CL2026
Wiki Dumps to Training Corpora: South Slavic Case
Mihailo Škorić, Cosimo Palma
This paper presents a pipeline designed to transform raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. T…
cs.CL2024
New Textual Corpora for Serbian Language Modeling
Mihailo Škorić, Nikola Janković
This paper will present textual corpora for Serbian (and Serbo-Croatian), usable for the training of large language models and publicly available at one of the several notable onli…
cs.CL2023
Text vectorization via transformer-based language models and n-gram perplexities
Mihailo Škorić
As the probability (and thus perplexity) of a text is calculated based on the product of the probabilities of individual tokens, it may happen that one unlikely token significantly…