3 papers
cs.LG2026
Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation
Callum Canavan, Aditya Shrivastava, Allison Qi +2
To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones…
cs.CL2025
LLM Optimization Unlocks Real-Time Pairwise Reranking
Jingyu Wu, Aditya Shrivastava, Jing Zhu +3
Efficiently reranking documents retrieved from information retrieval (IR) pipelines to enhance overall quality of Retrieval-Augmented Generation (RAG) system remains an important y…
cs.LG2025
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
Neel Jain, Aditya Shrivastava, Chenyang Zhu +6
A key component of building safe and reliable language models is enabling the models to appropriately refuse to follow certain instructions or answer certain questions. We may want…