Showing 2026Show all
2 papers · 1 filter
cs.CL2026
Value-Aware Numerical Representations for Transformer Language Models
Andreea Dutulescu, Stefan Ruseti, Mihai Dascalu
Transformer-based language models often achieve strong results on mathematical reasoning benchmarks while remaining fragile on basic numerical understanding and arithmetic operatio…
cs.CL2026
Training Language Models with homotokens Leads to Delayed Overfitting
Adrian Cosma, Stefan Ruseti, Emilian Radoi +1
Subword tokenization introduces a computational layer in language models where many distinct token sequences decode to the same surface form and preserve meaning, yet induce differ…