Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Recall Is Not Protection: Evaluating Safety Monitors Against Model Compliance
Sripad Karne
Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness labels, but…
cs.CL2026
How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings
Sripad Karne
Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary interface for understanding…
cs.CL2026
One Language, Two Scripts: Probing Script-Invariance in LLM Concept Representations
Sripad Karne
Do the features learned by Sparse Autoencoders (SAEs) represent abstract meaning, or are they tied to how text is written? We investigate this question using Serbian digraphia as a…