7 citations · 19 across the 12 of their papers we have counts for
12 papers
"Glue pizza and eat rocks" -- Exploiting Vulnerabilities in Retrieval-Augmented Generative Models
Zhen Tan, Chengshuai Zhao, Raha Moraffah +5
Retrieval-Augmented Generative (RAG) models enhance Large Language Models (LLMs) by integrating external knowledge bases, improving their performance in applications like fact-chec…
Cross-Platform Hate Speech Detection with Weakly Supervised Causal Disentanglement
Paras Sheth, Tharindu Kumarage, Raha Moraffah +2
Content moderation faces a challenging task as social media's ability to spread hate speech contrasts with its role in promoting global connectivity. With rapidly evolving slang an…
EAGLE: A Domain Generalization Framework for AI-generated Text Detection
Amrita Bhattacharjee, Raha Moraffah, Joshua Garland +1
With the advancement in capabilities of Large Language Models (LLMs), one major step in the responsible and safe use of such LLMs is to be able to detect text generated by these mo…
A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization
Tharindu Kumarage, Garima Agrawal, Paras Sheth +4
We have witnessed lately a rapid proliferation of advanced Large Language Models (LLMs) capable of generating high-quality text. While these LLMs have revolutionized text generatio…
Exploiting Class Probabilities for Black-box Sentence-level Attacks
Raha Moraffah, Huan Liu
Sentence-level attacks craft adversarial sentences that are synonymous with correctly-classified sentences but are misclassified by the text classifiers. Under the black-box settin…
Adversarial Text Purification: A Large Language Model Approach for Defense
Raha Moraffah, Shubh Khandelwal, Amrita Bhattacharjee +1
Adversarial purification is a defense mechanism for safeguarding classifiers against adversarial attacks without knowing the type of attacks or training of the classifier. These te…