5 papers
aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy
Fatih Deniz, Yazan Boshmaf, Dorde Popovic +1
The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or…
StructTransform: A Scalable Attack Surface for Safety-Aligned Large Language Models
Shehel Yoosuf, Temoor Ali, Ahmed Lekssays +2
In this work, we present a series of structure transformation attacks on LLM alignment, where we encode natural language intent using diverse syntax spaces, ranging from simple str…
aiXamine: Simplified LLM Safety and Security
Fatih Deniz, Dorde Popovic, Yazan Boshmaf +4
Evaluating Large Language Models (LLMs) for safety and security remains a complex task, often requiring users to navigate a fragmented landscape of ad hoc benchmarks, datasets, met…
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
Dorde Popovic, Amin Sadeghi, Ting Yu +2
Backdoor attacks are among the most effective, practical, and stealthy attacks in deep learning. In this paper, we consider a practical scenario where a developer obtains a deep mo…
MANTIS: Detection of Zero-Day Malicious Domains Leveraging Low Reputed Hosting Infrastructure
Fatih Deniz, Mohamed Nabeel, Ting Yu +1
Internet miscreants increasingly utilize short-lived disposable domains to launch various attacks. Existing detection mechanisms are either too late to catch such malicious domains…