activity
20192026
most citedNeural spell-checker: Beyond words with synthetic data generation

1 citations · 1 across the 1 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Challenges in Explaining Pretrained Clinical Text Classifiers

Kristian Miok, Matej Klemen, Blaz Škrlj +1

Explaining the predictions of neural models in clinical NLP remains a significant challenge, especially for complex tasks involving long, unstructured medical texts. While post-hoc…

cs.CL2026

Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages

Tjaša Arčon, Matej Klemen, Marko Robnik-Šikonja +1

LLMs are routinely evaluated on language use, yet their explicit knowledge about linguistic structure remains poorly understood. Existing linguistic benchmarks focus on narrow phen…

cs.CL2025

Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis

Matej Klemen, Tjaša Arčon, Luka Terčon +2

Empirical grammar research has become increasingly data-driven, but the systematic analysis of annotated corpora still requires substantial methodological and technical effort. We…

cs.CL20241 cited

Neural spell-checker: Beyond words with synthetic data generation

Matej Klemen, Martin Božič, Špela Arhar Holdt +1

Spell-checkers are valuable tools that enhance communication by identifying misspelled words in written texts. Recent improvements in deep learning, and in particular in large lang…

cs.CL2024

Code-mixed Sentiment and Hate-speech Prediction

Anjali Yadav, Tanya Garg, Matej Klemen +3

Code-mixed discourse combines multiple languages in a single text. It is commonly used in informal discourse in countries with several official languages, but also in many other co…