Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
Rohan Shravan
Large language models route every input through a learned embedding table of shape |V| x d_model, consuming hundreds of millions to billions of trainable parameters at frontier sca…
cs.CL2026
BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base
Rohan Shravan
We present BrahmicTokenizer-131K, a 131,072-vocabulary byte-level BPE tokenizer that closes the Brahmic compression gap at the 131K-vocabulary class while preserving the English, E…