1 citations · 1 across the 3 of their papers we have counts for
6 papers
Building a Strong Instruction Language Model for a Less-Resourced Language
Domen Vreš, Tjaša Arčon, Timotej Petrič +3
Large language models (LLMs) have become an essential tool for natural language processing and artificial intelligence in general. Current open-source models are primarily trained…
Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages
Tjaša Arčon, Matej Klemen, Marko Robnik-Šikonja +1
LLMs are routinely evaluated on language use, yet their explicit knowledge about linguistic structure remains poorly understood. Existing linguistic benchmarks focus on narrow phen…
Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis
Matej Klemen, Tjaša Arčon, Luka Terčon +2
Empirical grammar research has become increasingly data-driven, but the systematic analysis of annotated corpora still requires substantial methodological and technical effort. We…
Improving LLMs for Machine Translation Using Synthetic Preference Data
Dario Vajda, Domen Vreš, Marko Robnik-Šikonja
Large language models have emerged as effective machine translation systems. In this paper, we explore how a general instruction-tuned large language model can be improved for mach…
Neural spell-checker: Beyond words with synthetic data generation
Matej Klemen, Martin Božič, Špela Arhar Holdt +1
Spell-checkers are valuable tools that enhance communication by identifying misspelled words in written texts. Recent improvements in deep learning, and in particular in large lang…
Generative Model for Less-Resourced Language with 1 billion parameters
Domen Vreš, Martin Božič, Aljaž Potočnik +2
Large language models (LLMs) are a basic infrastructure for modern natural language processing. Many commercial and open-source LLMs exist for English, e.g., ChatGPT, Llama, Falcon…