From the 1 of 33 linked papers with an AI index.
4 papers · 1 filter
Mitigating Gradient Inversion Risks in Language Models via Token Obfuscation
Xinguo Feng, Zhongkui Ma, Zihan Wang +2
Training and fine-tuning large-scale language models largely benefit from collaborative learning, but the approach has been proven vulnerable to gradient inversion attacks (GIAs),…
Alert-ME: An Explainability-Driven Defense Against Adversarial Examples in Transformer-Based Text Classification
Bushra Sabir, Yansong Gao, Alsharif Abuadbba +1
Transformer-based text classifiers such as BERT, RoBERTa, T5, and GPT have shown strong performance in natural language processing tasks but remain vulnerable to adversarial exampl…
Adversarial Attacks Against Automated Fact-Checking: A Survey
Fanzhen Liu, Alsharif Abuadbba, Kristen Moore +5
In an era where misinformation spreads freely, fact-checking (FC) plays a crucial role in verifying claims and promoting reliable information. While automated fact-checking (AFC) h…
A Constraint-Enforcing Reward for Adversarial Attacks on Text Classifiers
Tom Roth, Inigo Jauregi Unanue, Alsharif Abuadbba +1
Text classifiers are vulnerable to adversarial examples -- correctly-classified examples that are deliberately transformed to be misclassified while satisfying acceptability constr…