4 papers
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats
Poojitha Thota, Yun Lei, Santhosh Thangaraj +2
Large language models (LLMs) are increasingly deployed in interactive applications, yet they remain vulnerable to adversarial interactions that induce harmful, deceptive, or policy…
Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning
Poojitha Thota, Shirin Nilizadeh
Training-time data poisoning during fine-tuning poses a significant threat to large language models (LLMs) deployed for abstractive text summarization, where small task-specific da…
FairDeFace: Evaluating the Fairness and Adversarial Robustness of Face Obfuscation Methods
Seyyed Mohammad Sadegh Moosavi Khorzooghi, Poojitha Thota, Mohit Singhal +3
The lack of a common platform and benchmark datasets for evaluating face obfuscation methods has been a challenge, with every method being tested using arbitrary experiments, datas…
Attacks against Abstractive Text Summarization Models through Lead Bias and Influence Functions
Poojitha Thota, Shirin Nilizadeh
Large Language Models have introduced novel opportunities for text comprehension and generation. Yet, they are vulnerable to adversarial perturbations and data poisoning attacks, p…