5 citations · 5 across the 4 of their papers we have counts for
3 papers · 1 filter
aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy
Fatih Deniz, Yazan Boshmaf, Dorde Popovic +1
The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or…
aiXamine: Simplified LLM Safety and Security
Fatih Deniz, Dorde Popovic, Yazan Boshmaf +4
Evaluating Large Language Models (LLMs) for safety and security remains a complex task, often requiring users to navigate a fragmented landscape of ad hoc benchmarks, datasets, met…
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
Dorde Popovic, Amin Sadeghi, Ting Yu +2
Backdoor attacks are among the most effective, practical, and stealthy attacks in deep learning. In this paper, we consider a practical scenario where a developer obtains a deep mo…