Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Unifying Linear-Time Attention via Latent Probabilistic Modelling
Rares Dolga, Lucas Maystre, Marius Cobzarenco +1
Transformers have achieved state-of-the-art results across a range of domains, but their quadratic attention mechanism poses significant challenges for long-sequence modelling. Rec…
cs.CL2025
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
Rares Dolga, Lucas Maystre, Tudor Berariu +1
Subword tokenization methods like Byte Pair Encoding (BPE) are widely used in large language models due to their balance of vocabulary compactness and representational power. Howev…