3 papers
cs.CR2026
GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis
Subrat Kishore Dutta, Yuelin Xu, Piyush Pant +1
Recent work has shown that RLHF is highly susceptible to backdoor attacks. However, existing methods often rely on rare tokens or fixed triggers, limiting their impact in realistic…
cs.CV2025
Generalizable Targeted Data Poisoning against Varying Physical Objects
Zhizhen Chen, Zhengyu Zhao, Subrat Kishore Dutta +3
Targeted data poisoning (TDP) aims to compromise the model's prediction on a specific (test) target by perturbing a small subset of training data. Existing work on TDP has focused…
cs.CV2025
IAP: Invisible Adversarial Patch Attack through Perceptibility-Aware Localization and Perturbation Optimization
Subrat Kishore Dutta, Xiao Zhang
Despite modifying only a small localized input region, adversarial patches can drastically change the prediction of computer vision models. However, prior methods either cannot per…