6 papers
The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation
Anton M. Alekseev, Sergey I. Nikolenko
Dungan, a Sinitic language of Central Asia written in a Cyrillic-based script, is described in detail in the grammatical literature, yet the quantitative properties of its morpholo…
KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding
Timur Turatali, Aida Turdubaeva, Rustem Izmailov +2
Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cu…
Query2Diagram: Answering Developer Queries with UML Diagrams
Oleg Baryshnikov, Anton M. Alekseev, Sergey I. Nikolenko
Software documentation frequently becomes outdated or fails to exist entirely, yet developers need focused views of their codebase to understand complex systems. While automated re…
Human-Annotated NER Dataset for the Kyrgyz Language
Timur Turatali, Anton Alekseev, Gulira Jumalieva +2
We introduce KyrgyzNER, the first manually annotated named entity recognition dataset for the Kyrgyz language. Comprising 1,499 news articles from the 24.KG news portal, the datase…
HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings
Anton Alekseev, Gulnara Kabaeva
One of the key tasks in modern applied computational linguistics is constructing word vector representations (word embeddings), which are widely used to address natural language pr…
KyrgyzNLP: Challenges, Progress, and Future
Anton Alekseev, Timur Turatali
Large language models (LLMs) have excelled in numerous benchmarks, advancing AI applications in both linguistic and non-linguistic tasks. However, this has primarily benefited well…