Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Regress, Don't Guess -- A Regression-like Loss on Number Tokens for Language Models
Jonas Zausinger, Lars Pennig, Anamarija Kozina +13
While language models have exceptional capabilities at text generation, they lack a natural inductive bias for emitting numbers and thus struggle in tasks involving quantitative re…
cs.CL2025
Multiscale Byte Language Models -- A Hierarchical Architecture for Causal Million-Length Sequence Modeling
Eric Egli, Matteo Manica, Jannis Born
Bytes form the basis of the digital world and thus are a promising building block for multimodal foundation models. Recently, Byte Language Models (BLMs) have emerged to overcome t…