6 papers
The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation
Anton M. Alekseev, Sergey I. Nikolenko
Dungan, a Sinitic language of Central Asia written in a Cyrillic-based script, is described in detail in the grammatical literature, yet the quantitative properties of its morpholo…
KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding
Timur Turatali, Aida Turdubaeva, Rustem Izmailov +2
Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cu…
Human-Annotated NER Dataset for the Kyrgyz Language
Timur Turatali, Anton Alekseev, Gulira Jumalieva +2
We introduce KyrgyzNER, the first manually annotated named entity recognition dataset for the Kyrgyz language. Comprising 1,499 news articles from the 24.KG news portal, the datase…
Syntactic Transfer to Kyrgyz Using the Treebank Translation Method
Anton Alekseev, Alina Tillabaeva, Gulnara Dzh. Kabaeva +1
The Kyrgyz language, as a low-resource language, requires significant effort to create high-quality syntactic corpora. This study proposes an approach to simplify the development p…
DFT: A Universal Quantum Chemistry Dataset of Drug-Like Molecules and a Benchmark for Neural Network Potentials
Kuzma Khrabrov, Anton Ber, Artem Tsypin +10
Methods of computational quantum chemistry provide accurate approximations of molecular properties crucial for computer-aided drug discovery and other areas of chemical science. Ho…
Neural Click Models for Recommender Systems
Mikhail Shirokikh, Ilya Shenbin, Anton Alekseev +4
We develop and evaluate neural architectures to model the user behavior in recommender systems (RS) inspired by click models for Web search but going beyond standard click models.…