papers

Publications (11)

cs.CL2025

Towards Generating Automatic Anaphora Annotations

Dima Taji, Daniel Zeman

Training models that can perform well on various NLP tasks require large amounts of data, and this becomes more apparent with nuanced tasks such as anaphora and conference resoluti…

cs.CL2022

Findings of the Shared Task on Multilingual Coreference Resolution

Zdeněk Žabokrtský, Miloslav Konopík, Anna Nedoluzhko +7

This paper presents an overview of the shared task on multilingual coreference resolution associated with the CRAC 2022 workshop. Shared task participants were supposed to develop…

cs.CL2020

Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection

Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter +6

Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework.…

cs.CL2024

Findings of the Third Shared Task on Multilingual Coreference Resolution

Michal Novák, Barbora Dohnalová, Miloslav Konopík +7

The paper presents an overview of the third edition of the shared task on multilingual coreference resolution, held as part of the CRAC 2024 workshop. Similarly to the previous two…

cs.CL2023

A Unified Taxonomy of Deep Syntactic Relations

Kira Droganova, Daniel Zeman

This paper analyzes multiple deep-syntactic frameworks with the goal of creating a proposal for a set of universal semantic role labels. The proposal examines various theoretic lin…

cs.CL2000

Automatic Extraction of Subcategorization Frames for Czech

Anoop Sarkar, Daniel Zeman

We present some novel machine learning techniques for the identification of subcategorization information for verbs in Czech. We compare three different statistical techniques appl…

cs.CL2026

Meet UD_Czech-PDTC: A Large and Genre-Rich Treebank in Universal Dependencies

Marie Mikulová, Barbora Štěpánková, Daniel Zeman +3

Czech has been part of Universal Dependencies since its first release in 2015. It has also been one of the best represented languages, with the Prague Dependency Treebank being ord…

cs.CL2026

Word Alignment-Based Evaluation of Uniform Meaning Representations

Daniel Zeman, Federica Gamba

Comparison and evaluation of graph-based representations of sentence meaning is a challenge because competing representations of the same sentence may have different number of node…

cs.CL2026

Findings of the Fifth Shared Task on Multilingual Coreference Resolution: Expanding Datasets for Long-Range Entities

Michal Novák, Miloslav Konopík, Anna Nedoluzhko +6

This paper describes the fifth edition of the Shared Task on Multilingual Coreference Resolution, held in conjunction with the CODI-CRAC 2026 workshop. Building on previous iterati…

cs.CL2025

Findings of the Fourth Shared Task on Multilingual Coreference Resolution: Can LLMs Dethrone Traditional Approaches?

Michal Novák, Miloslav Konopík, Anna Nedoluzhko +6

The paper presents an overview of the fourth edition of the Shared Task on Multilingual Coreference Resolution, organized as part of the CODI-CRAC 2025 workshop. As in the previous…

cs.CL2020

Predicting Typological Features in WALS using Language Embeddings and Conditional Probabilities: ÚFAL Submission to the SIGTYP 2020 Shared Task

Martin Vastl, Daniel Zeman, Rudolf Rosa

We present our submission to the SIGTYP 2020 Shared Task on the prediction of typological features. We submit a constrained system, predicting typological features only based on th…