3 papers
cs.CL2026
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
Sachin Kumar
Linear probes trained on LLM activations are increasingly proposed as deception-detection metrics, yet report AUROC exceeding 0.96 on clean benchmarks while collapsing under distri…
cs.CL2026
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
Sachin Kumar
Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior when triggered by specific patt…
cs.CL2026
Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models
Sachin Kumar
Can small language models achieve strong tool-use performance without complex adaptation mechanisms? This paper investigates this question through Meta-Tool, a controlled empirical…