6 papers
Anatomy of a Sound Neural Reasoner: One-Shot Amortization, First-Pass Poisoning, and Search Inertness in Clue-Rich Completion
Aleksey Komissarov
Neural solvers are built to deduce, branch, and revise intermediate states. The Lattice Deduction Transformer (LDT) appears to do exactly that. In clue-rich Sudoku, it does not: on…
Compressed code: the hidden effects of quantization and distillation on programming tokens
Viacheslav Siniaev, Iaroslav Chelombitko, Aleksey Komissarov
Large Language Models (LLMs) have demonstrated exceptional code generation capabilities, yet their token-level mechanisms remain underexplored, particularly in compressed models. T…
Subword-Based Comparative Linguistics across 242 Languages Using Wikipedia Glottosets
Iaroslav Chelombitko, Mika Hämäläinen, Aleksey Komissarov
We present a large-scale comparative study of 242 Latin and Cyrillic-script languages using subword-based methodologies. By constructing 'glottosets' from Wikipedia lexicons, we in…
SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
Iaroslav Chelombitko, Ekaterina Chelombitko, Aleksey Komissarov
The quality of subword tokenization is critical for Large Language Models, yet evaluating tokenizers for morphologically rich Uralic languages is hampered by the lack of clean morp…
When repeats drive the vocabulary: a Byte-Pair Encoding analysis of T2T primate genomes
Marina Popova, Iaroslav Chelombitko, Aleksey Komissarov
The emergence of telomere-to-telomere (T2T) genome assemblies has opened new avenues for comparative genomics, yet effective tokenization strategies for genomic sequences remain un…
Qtok: A Comprehensive Framework for Evaluating Multilingual Tokenizer Quality in Large Language Models
Iaroslav Chelombitko, Egor Safronov, Aleksey Komissarov
In the development of Large Language Models (LLMs), considerable attention has been given to the quality of training datasets. However, the role of tokenizers in the LLM training p…