activity
20222024
most citedDomain Generalization -- A Causal Perspective

7 citations · 19 across the 12 of their papers we have counts for

collaborators

12 papers

cs.CR2024

"Glue pizza and eat rocks" -- Exploiting Vulnerabilities in Retrieval-Augmented Generative Models

Zhen Tan, Chengshuai Zhao, Raha Moraffah +5

Retrieval-Augmented Generative (RAG) models enhance Large Language Models (LLMs) by integrating external knowledge bases, improving their performance in applications like fact-chec…

cs.LG20241 cited

Cross-Platform Hate Speech Detection with Weakly Supervised Causal Disentanglement

Paras Sheth, Tharindu Kumarage, Raha Moraffah +2

Content moderation faces a challenging task as social media's ability to spread hate speech contrasts with its role in promoting global connectivity. With rapidly evolving slang an…

cs.CL20243 cited

EAGLE: A Domain Generalization Framework for AI-generated Text Detection

Amrita Bhattacharjee, Raha Moraffah, Joshua Garland +1

With the advancement in capabilities of Large Language Models (LLMs), one major step in the responsible and safe use of such LLMs is to be able to detect text generated by these mo…

cs.CL20243 cited

A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization

Tharindu Kumarage, Garima Agrawal, Paras Sheth +4

We have witnessed lately a rapid proliferation of advanced Large Language Models (LLMs) capable of generating high-quality text. While these LLMs have revolutionized text generatio…

cs.CL2024

Exploiting Class Probabilities for Black-box Sentence-level Attacks

Raha Moraffah, Huan Liu

Sentence-level attacks craft adversarial sentences that are synonymous with correctly-classified sentences but are misclassified by the text classifiers. Under the black-box settin…

cs.CR2024

Adversarial Text Purification: A Large Language Model Approach for Defense

Raha Moraffah, Shubh Khandelwal, Amrita Bhattacharjee +1

Adversarial purification is a defense mechanism for safeguarding classifiers against adversarial attacks without knowing the type of attacks or training of the classifier. These te…