4 papers
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
Phillip Howard, Xin Su, Allen Roush +2
High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling techniques have helped mitigate drawbacks…
Riemannian-Manifold Steering: Geometry-Aware Generative Autoencoders for Label-Free Steering
Narmeen Oozeer, Shivam Raval, Philip Quirke +4
Steering a language model - intervening on its internal activations to change downstream behaviour - has recently expanded beyond linear interpolation to nonlinear methods such as…
INDIC DIALECT: A Multi Task Benchmark to Evaluate and Translate in Indian Language Dialects
Tarun Sharma, Manikandan Ravikiran, Sourava Kumar Behera +3
Recent NLP advances focus primarily on standardized languages, leaving most low-resource dialects under-served especially in Indian scenarios. In India, the issue is particularly i…
Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement
Karrtik Iyer, Manikandan Ravikiran, Prasanna Pendse +1
Automated grading systems can efficiently score short-answer responses, yet they often fail to indicate when a grading decision is uncertain or potentially contentious. We introduc…