4 papers
Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations
Alizishaan Khatri
LLM-based code generation is now embedded in mission-critical pipelines, but defenses against vulnerable output remain post-hoc -- static analyzers, fine-tuned classifiers, or an L…
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families
Alizishaan Khatri, Dun Li Chan
Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-8B) detect harmful prompts at…
Latent Space Probing for Adult Content Detection in Video Generative Models
Alizishaan Khatri, Chiquita Prabhu
The rapid proliferation of AI-powered video generation systems has introduced significant challenges in content moderation, particularly with respect to adult and sexually explicit…
Preventing overfitting in deep learning using differential privacy
Alizishaan Anwar Hussein Khatri
The use of Deep Neural Network based systems in the real world is growing. They have achieved state-of-the-art performance on many image, speech and text datasets. They have been s…