activity
20242026
most citedLLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification

13 citations · 15 across the 13 of their papers we have counts for

collaborators
Showing 2026Show all

5 papers · 1 filter

cs.CL2026

Opinionated, Hesitant and Stressed: Three Studies of How Politicians Speak in Four Slavic Parliaments

Ivan Porupski, Nikola Ljubešić

We present three large-scale studies of spoken parliamentary speech across four Slavic languages (Croatian, Czech, Polish, Serbian), drawing on over 6,000 hours from the ParlaSpeec…

cs.CL2026

Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments

Ivan Porupski, Branimir Dropuljić, Nikola Ljubešić

Filled pauses (FPs) are a universal feature of spontaneous speech, yet most studies rely on small, single-language corpora, limiting the generalisability of their findings. We anal…

cs.CL2026

The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora

Taja Kuzman Pungeršek, Peter Rupnik, Vít Suchomel +1

Crawling national top-level domains has proven to be highly effective for collecting texts in less-resourced languages. This approach has been recently used for South Slavic langua…

eess.AS2026

Mići Princ -- A Little Boy Teaching Speech Technologies the Chakavian Dialect

Nikola Ljubešić, Peter Rupnik, Tea Perinčić

This paper documents our efforts in releasing the printed and audio book of the translation of the famous novel The Little Prince into the Chakavian dialect, as a computer-readable…

cs.CL2026

Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification

Taja Kuzman Pungeršek, Peter Rupnik, Daniela Širinić +1

This paper introduces ParlaCAP, a large-scale dataset for analyzing parliamentary agenda setting across Europe, and proposes a cost-effective method for building domain-specific po…