8 papers
Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark
Terra Blevins, Stephen Mayhew, Marek Å uppa +11
While multilingual language models promise to bring the benefits of LLMs to speakers of many languages, gold-standard evaluation benchmarks in most languages to interrogate these a…
Coconstructions in spoken data: UD annotation guidelines and first results
Ludovica Pannitto, Sylvain Kahane, Kaja Dobrovoljc +4
The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns - including collaborative coconstructions proper, wh-question answers, and backch…
A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding
Dilara TorunoÄlu-Selamet, Dogukan Arslan, Rodrigo Wilkens +75
Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challen…
Counting trees: A treebank-driven exploration of syntactic variation in speech and writing across languages
Kaja Dobrovoljc
This paper presents a novel treebank-driven approach to comparing syntactic structures in speech and writing using dependency-parsed corpora. Adopting a fully inductive, bottom-up…
Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages
TjaÅ¡a ArÄon, Matej Klemen, Marko Robnik-Å ikonja +1
LLMs are routinely evaluated on language use, yet their explicit knowledge about linguistic structure remains poorly understood. Existing linguistic benchmarks focus on narrow phen…
Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis
Matej Klemen, TjaÅ¡a ArÄon, Luka TerÄon +2
Empirical grammar research has become increasingly data-driven, but the systematic analysis of annotated corpora still requires substantial methodological and technical effort. We…