2 papers
cs.CL2025
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
Vladislav Stankov, Matyáš Kopp, Ondřej Bojar
We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the…
cs.CL2024
Multilingual Power and Ideology Identification in the Parliament: a Reference Dataset and Simple Baselines
Çağrı Çöltekin, Matyáš Kopp, Katja Meden +3
We introduce a dataset on political orientation and power position identification. The dataset is derived from ParlaMint, a set of comparable corpora of transcribed parliamentary s…