3 papers
cs.CL2025
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
Mykyta Syromiatnikov, Victoria Ruvinskaya
Evaluating the real capabilities of large language models in low-resource languages still represents a challenge, as many existing benchmarks focus on widespread tasks translated f…
cs.CL2025
Empowering Smaller Models: Tuning LLaMA and Gemma with Chain-of-Thought for Ukrainian Exam Tasks
Mykyta Syromiatnikov, Victoria Ruvinskaya, Nataliia Komleva
Leading large language models have demonstrated impressive capabilities in reasoning-intensive tasks, such as standardized educational testing. However, they often require extensiv…
cs.CL2025
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
Mykyta Syromiatnikov, Victoria Ruvinskaya, Anastasiya Troynina
As the usage of large language models for problems outside of simple text understanding or generation increases, assessing their abilities and limitations becomes crucial. While si…