most citedNBA: defensive distillation for backdoor removal via neural behavior alignment

6 citations · 12 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2025

PRJ: Perception-Retrieval-Judgement for Generated Images

Qiang Fu, Zonglei Jing, Zonghao Ying +1

The rapid progress of generative AI has enabled remarkable creative capabilities, yet it also raises urgent concerns regarding the safety of AI-generated visual content in real-wor…

cs.CR2024

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity

Zhengmin Yu, Jiutian Zeng, Siyi Chen +7

Over the past year, there has been a notable rise in the use of large language models (LLMs) for academic research and industrial practices within the cybersecurity field. However,…

cs.CR20243 cited

Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks

Zonghao Ying, Aishan Liu, Xianglong Liu +1

The recent release of GPT-4o has garnered widespread attention due to its powerful general capabilities. While its impressive performance is widely acknowledged, its safety aspects…

cs.CV20241 cited

Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Zonghao Ying, Aishan Liu, Tianyuan Zhang +4

In the realm of large vision language models (LVLMs), jailbreak attacks serve as a red-teaming approach to bypass guardrails and uncover safety implications. Existing jailbreaks pr…

cs.CR20242 cited

DLP: towards active defense against backdoor attacks with decoupled learning process

Zonghao Ying, Bin Wu

Deep learning models are well known to be susceptible to backdoor attack, where the attacker only needs to provide a tampered dataset on which the triggers are injected. Models tra…

cs.CR20246 cited

NBA: defensive distillation for backdoor removal via neural behavior alignment

Zonghao Ying, Bin Wu

Recently, deep neural networks have been shown to be vulnerable to backdoor attacks. A backdoor is inserted into neural networks via this attack paradigm, thus compromising the int…