Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems
Shwetha Singaravelu, Gayathri Muruganantham, Lakshmi Rajendran +1
Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language gen…
cs.CL2026
BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis
Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran +1
Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentatio…