3 papers
cs.LG2026
Adversarial Robustness of Activation Steering in Large Language Models
Kien Le, Thai Le
Activation steering has become a popular training-free method to control LLM behavior by injecting precomputed direction vectors into the model's residual stream at inference time.…
cs.CL2026
SARA: Stress Test Reasoning in Audio Deepfake Detection
Binh Nguyen, Charles Fleming, Thai Le
Audio Language Models (ALMs) offer a promising shift towards explainable audio deepfake detections (ADD), moving beyond \textit{black-box} classifiers by providing transparency to…
cs.LG2025
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
Binh Nguyen, Shuji Shi, Ryan Ofman +1
Recent advances in text-to-speech technologies have enabled realistic voice generation, fueling audio-based deepfake attacks such as fraud and impersonation. While audio anti-spoof…