3 papers
cs.AI2026
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
Caleb DeLeeuw
Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model refuses, is that refusal st…
cs.CY2025
The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind
Caleb DeLeeuw, Gaurav Chawla, Aniket Sharma +1
We investigate strategic deception in large language models using two complementary testbeds: Secret Agenda (across 38 models) and Insider Trading compliance (via SAE architectures…
cs.CL2025
Real-World En Call Center Transcripts Dataset with PII Redaction
Ha Dao, Gaurav Chawla, Raghu Banda +1
We introduce CallCenterEN, a large-scale (91,706 conversations, corresponding to 10448 audio hours), real-world English call center transcript dataset designed to support research…