6 citations · 20 across the 30 of their papers we have counts for
4 papers · 1 filter
Probing Subphonemes in Morphology Models
Gal Astrach, Yuval Pinter
Transformers have achieved state-of-the-art performance in morphological inflection tasks, yet their ability to generalize across languages and morphological rules remains limited.…
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
Craig W. Schmidt, Varshini Reddy, Chris Tanner +1
Pre-tokenization, the initial step in many modern tokenization pipelines, segments text into smaller units called pretokens, typically splitting on whitespace and punctuation. Whil…
Token-Level Privacy in Large Language Models
Re'em Harel, Niv Gilboa, Yuval Pinter
The use of language models as remote services requires transmitting private information to external providers, raising significant privacy concerns. This process not only risks exp…
How Much is Enough? The Diminishing Returns of Tokenization Training Data
Varshini Reddy, Craig W. Schmidt, Yuval Pinter +1
Tokenization, a crucial initial step in natural language processing, is governed by several key parameters, such as the tokenization algorithm, vocabulary size, pre-tokenization st…