Showing cs.CRShow all
3 papers · 1 filter
cs.CR2026
Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense
Subhadip Mitra
Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no defense, static steering, CAS…
cs.CR2026
Cross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMs
Subhadip Mitra
Safety alignment in LLMs does not improve monotonically across model generations. Studying four generations of Google's Gemma family (7B-31B) with quality-diversity evolution (MAP-…
cs.CR2026
Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety
Subhadip Mitra
Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not scale, LLM-as-attacker methods exhibit mode collapse, and gradient-based approa…