◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Mihailo Škorić

3 papers hereh-index 592 citations23 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author1
  • first author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL3

identity via Semantic Scholar / OpenAlex

activity
20232026
collaborators

3 papers

cs.CL2026

Wiki Dumps to Training Corpora: South Slavic Case

Mihailo Škorić, Cosimo Palma

This paper presents a pipeline designed to transform raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. T…

cs.CL2024

New Textual Corpora for Serbian Language Modeling

Mihailo Škorić, Nikola Janković

This paper will present textual corpora for Serbian (and Serbo-Croatian), usable for the training of large language models and publicly available at one of the several notable onli…

cs.CL2023

Text vectorization via transformer-based language models and n-gram perplexities

Mihailo Škorić

As the probability (and thus perplexity) of a text is calculated based on the product of the probabilities of individual tokens, it may happen that one unlikely token significantly…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.