fine-tuning 1instruction tuning 1multilingual language models 1token fragmentation 1tokenization robustness 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
Language Models are not Equally Robust to Non-Canonical Tokenization across Languages
Poulami Ghosh, Preethi Jyothi
The paper investigates how language models handle non-canonical tokenizations across 27 languages, finding that robustness varies by language and token fragmentation, and shows tha…
cs.CL2024
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
Poulami Ghosh, Raj Dabre, Pushpak Bhattacharyya
Pre-trained language models (PLMs) are known to be susceptible to perturbations to the input text, but existing works do not explicitly focus on linguistically grounded attacks, wh…