2 papers
cs.CL2026
Does the LM Head Create a Harmful Gradient Bottleneck? A Causal Test
Anand Murugan
The language-model head maps a hidden state of width D to a vocabulary of size V, so its transpose can return at most D independent directions to the Transformer. Godey and Artzi a…
cs.CL2026
Morphology Aware Reversible Semantic Tokenization and Hierarchical Word Composition for Tamil Language Models
Anand Murugan
Statistical subword tokenizers can process arbitrary text, but their units need not align with lexical or grammatical structure. This is especially important for Tamil, where a wri…