3 papers
cs.CL2026
KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding
Timur Turatali, Aida Turdubaeva, Rustem Izmailov +2
Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cu…
cs.CL2025
Human-Annotated NER Dataset for the Kyrgyz Language
Timur Turatali, Anton Alekseev, Gulira Jumalieva +2
We introduce KyrgyzNER, the first manually annotated named entity recognition dataset for the Kyrgyz language. Comprising 1,499 news articles from the 24.KG news portal, the datase…
cs.CL2024
KyrgyzNLP: Challenges, Progress, and Future
Anton Alekseev, Timur Turatali
Large language models (LLMs) have excelled in numerous benchmarks, advancing AI applications in both linguistic and non-linguistic tasks. However, this has primarily benefited well…