2 papers
cs.CL2025
The Tokenization Bottleneck: How Vocabulary Extension Improves Chemistry Representation Learning in Pretrained Language Models
Prathamesh Kalamkar, Ned Letcher, Meissane Chami +3
The application of large language models (LLMs) to chemistry is frequently hampered by a "tokenization bottleneck", where tokenizers tuned on general-domain text tend to fragment c…
cs.AI2025
Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement
Karrtik Iyer, Manikandan Ravikiran, Prasanna Pendse +1
Automated grading systems can efficiently score short-answer responses, yet they often fail to indicate when a grading decision is uncertain or potentially contentious. We introduc…