3 papers
cs.CL2025
Human-Annotated NER Dataset for the Kyrgyz Language
Timur Turatali, Anton Alekseev, Gulira Jumalieva +2
We introduce KyrgyzNER, the first manually annotated named entity recognition dataset for the Kyrgyz language. Comprising 1,499 news articles from the 24.KG news portal, the datase…
cs.CL2024
HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings
Anton Alekseev, Gulnara Kabaeva
One of the key tasks in modern applied computational linguistics is constructing word vector representations (word embeddings), which are widely used to address natural language pr…
cs.CL2024
KyrgyzNLP: Challenges, Progress, and Future
Anton Alekseev, Timur Turatali
Large language models (LLMs) have excelled in numerous benchmarks, advancing AI applications in both linguistic and non-linguistic tasks. However, this has primarily benefited well…