activity
20192025
most citedMetric Learning for Adversarial Robustness

58 citations · 90 across the 19 of their papers we have counts for

collaborators
Showing 2024Show all

7 papers · 1 filter

cs.CL2024

Diversity Helps Jailbreak Large Language Models

Weiliang Zhao, Daniel Ben-Levi, Wei Hao +2

We have uncovered a powerful jailbreak technique that leverages large language models' ability to diverge from prior context, enabling them to bypass safety constraints and generat…

cs.SD20243 cited

I Can Hear You: Selective Robust Training for Deepfake Audio Detection

Zirui Zhang, Wei Hao, Aroon Sankoh +4

Recent advances in AI-generated voices have intensified the challenge of detecting deepfake audio, posing risks for scams and the spread of disinformation. To tackle this issue, we…

cs.CL2024

SPIN: Self-Supervised Prompt INjection

Leon Zhou, Junfeng Yang, Chengzhi Mao

Large Language Models (LLMs) are increasingly used in a variety of important applications, yet their safety and reliability remain as major concerns. Various adversarial and jailbr…

cs.CL20241 cited

RAFT: Realistic Attacks to Fool Text Detectors

James Wang, Ran Li, Junfeng Yang +1

Large language models (LLMs) have exhibited remarkable fluency across various tasks. However, their unethical applications, such as disseminating disinformation, have become a grow…

cs.CL20241 cited

Learning to Rewrite: Generalized LLM-Generated Text Detection

Ran Li, Wei Hao, Weiliang Zhao +2

Large language models (LLMs) present significant risks when used to generate non-factual content and spread disinformation at scale. Detecting such LLM-generated content is crucial…

cs.CV20241 cited

Turns Out I'm Not Real: Towards Robust Detection of AI-Generated Videos

Qingyuan Liu, Pengyuan Shi, Yun-Yun Tsai +2

The impressive achievements of generative models in creating high-quality videos have raised concerns about digital integrity and privacy vulnerabilities. Recent works to combat De…