2 papers
cs.LG2026
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
Shravan Doda
Final-token safety probes monitor a single hidden state after prompt prefill, but jailbreak prompts can contain probe-visible unsafe evidence distributed across earlier user-token…
cs.LG2025
VERBA: Verbalizing Model Differences Using Large Language Models
Shravan Doda, Shashidhar Reddy Javaji, Zining Zhu
In the current machine learning landscape, we face a "model lake" phenomenon: Given a task, there is a proliferation of trained models with similar performances despite different b…