1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2025
UNIDOOR: A Universal Framework for Action-Level Backdoor Attacks in Deep Reinforcement Learning
Oubo Ma, Linkang Du, Yang Dai +4
Deep reinforcement learning (DRL) is widely applied to safety-critical decision-making scenarios. However, DRL is vulnerable to backdoor attacks, especially action-level backdoors,…
cs.CR2024
CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP Models
Rui Zeng, Xi Chen, Yuwen Pu +3
Backdoors can be injected into NLP models to induce misbehavior when the input text contains a specific feature, known as a trigger, which the attacker secretly selects. Unlike fix…
cs.AI2022★ 1 cited
"Is your explanation stable?": A Robustness Evaluation Framework for Feature Attribution
Yuyou Gan, Yuhao Mao, Xuhong Zhang +5
Understanding the decision process of neural networks is hard. One vital method for explanation is to attribute its decision to pivotal features. Although many algorithms are propo…