4 papers · 1 filter
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
Sachin Kumar
Linear probes trained on LLM activations are increasingly proposed as deception-detection metrics, yet report AUROC exceeding 0.96 on clean benchmarks while collapsing under distri…
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
Sachin Kumar
Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior when triggered by specific patt…
Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models
Sachin Kumar
Can small language models achieve strong tool-use performance without complex adaptation mechanisms? This paper investigates this question through Meta-Tool, a controlled empirical…
Overriding Safety protections of Open-source Models
Sachin Kumar
LLMs(Large Language Models) nowadays have widespread adoption as a tool for solving issues across various domain/tasks. These models since are susceptible to produce harmful or tox…