3 papers
cs.CL2025
Broken Words, Broken Performance: Effect of Tokenization on Performance of LLMs
Sachin Pawar, Manoj Apte, Kshitij Jadhav +2
Tokenization is the first step in training any Large Language Model (LLM), where the text is split into a sequence of tokens as per the model's fixed vocabulary. This tokenization…
cs.CL2025
Explainable Statute Prediction via Attention-based Model and LLM Prompting
Sachin Pawar, Girish Keshav Palshikar, Anindita Sinha Banerjee +2
In this paper, we explore the problem of automatic statute prediction where for a given case description, a subset of relevant statutes are to be predicted. Here, the term "statute…
cs.CL2025
Matching Tasks with Industry Groups for Augmenting Commonsense Knowledge
Rituraj Singh, Sachin Pawar, Girish Palshikar
Commonsense knowledge bases (KB) are a source of specialized knowledge that is widely used to improve machine learning applications. However, even for a large KB such as ConceptNet…