10 papers
Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction
Yuqiao Xu, Osama Zafar, Alexander Nemecek +1
Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can protect users during ongoing interactions…
Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking
Alexander Nemecek, Osama Zafar, Yuqiao Xu +2
Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure for content provenance. Yet a…
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
Osama Zafar, Alexander Nemecek, Yiqian Zhang +5
Standard PII filters often miss contextual data leakage in RAG systems, such as non-regulated attribute clusters that collectively identify individuals. We introduce a Privacy Poli…
Reliability-Gated Source Anchoring for Continual Test-Time Adaptation
Vikash Singh, Debargha Ganguly, Weicong Chen +8
Continual test-time adaptation (CTTA) updates a pretrained model online on an unlabeled, non-stationary stream while anchoring it to a frozen source checkpoint. This anchor is usef…
The End of Trust: How Agentic AI Breaks Security Assumptions
Osama Zafar, Alexander Nemecek, Erman Ayday
For decades, the security of digital interaction has rested on an unacknowledged economic constraint. Attackers faced a tradeoff between the fidelity of a deception and the scale a…
Neighborhood Blending: A Lightweight Inference-Time Defense Against Membership Inference Attacks
Osama Zafar, Shaojie Zhan, Tianxi Ji +1
In recent years, the widespread adoption of Machine Learning as a Service (MLaaS), particularly in sensitive environments, has raised considerable privacy concerns. Of particular i…