3 papers
cs.CL2026
Language Models are not Equally Robust to Non-Canonical Tokenization across Languages
Poulami Ghosh, Preethi Jyothi
Despite the existence of exponentially many valid tokenizations for a given string, language models operate on a single canonical sequence deterministically produced by the tokeniz…
cs.CL2024
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
Poulami Ghosh, Raj Dabre, Pushpak Bhattacharyya
Pre-trained language models (PLMs) are known to be susceptible to perturbations to the input text, but existing works do not explicitly focus on linguistically grounded attacks, wh…
cs.CL2024
A Morphology-Based Investigation of Positional Encodings
Poulami Ghosh, Shikhar Vashishth, Raj Dabre +1
Contemporary deep learning models effectively handle languages with diverse morphology despite not being directly integrated into them. Morphology and word order are closely linked…