activity
20242026
collaborators

6 papers

cs.CL2026

Building a Strong Instruction Language Model for a Less-Resourced Language

Domen Vreš, Tjaša Arčon, Timotej Petrič +3

Large language models (LLMs) have become an essential tool for natural language processing and artificial intelligence in general. Current open-source models are primarily trained…

cs.CL2026

Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages

Tjaša Arčon, Matej Klemen, Marko Robnik-Šikonja +1

LLMs are routinely evaluated on language use, yet their explicit knowledge about linguistic structure remains poorly understood. Existing linguistic benchmarks focus on narrow phen…

cs.CL2025

Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis

Matej Klemen, Tjaša Arčon, Luka Terčon +2

Empirical grammar research has become increasingly data-driven, but the systematic analysis of annotated corpora still requires substantial methodological and technical effort. We…

cs.CL2025

Improving LLMs for Machine Translation Using Synthetic Preference Data

Dario Vajda, Domen Vreš, Marko Robnik-Šikonja

Large language models have emerged as effective machine translation systems. In this paper, we explore how a general instruction-tuned large language model can be improved for mach…

cs.CL2024

Neural spell-checker: Beyond words with synthetic data generation

Matej Klemen, Martin Božič, Špela Arhar Holdt +1

Spell-checkers are valuable tools that enhance communication by identifying misspelled words in written texts. Recent improvements in deep learning, and in particular in large lang…

cs.CL2024

Generative Model for Less-Resourced Language with 1 billion parameters

Domen Vreš, Martin Božič, Aljaž Potočnik +2

Large language models (LLMs) are a basic infrastructure for modern natural language processing. Many commercial and open-source LLMs exist for English, e.g., ChatGPT, Llama, Falcon…