4 papers
DMRL: Data- and Model-aware Reward Learning for Data Extraction
Zhiqiang Wang, Ruoxi Cheng
Large language models (LLMs) are inherently vulnerable to unintended privacy breaches. Consequently, systematic red-teaming research is essential for developing robust defense mech…
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
Ruoxi Cheng, Yizhong Ding, Shuirong Cao +6
Understanding the vulnerabilities of Large Vision Language Models (LVLMs) to jailbreak attacks is essential for their responsible real-world deployment. Most previous work requires…
Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining
Ruoxi Cheng, Yizhong Ding, Shuirong Cao +2
Audio can disclose PII, particularly when combined with related text data. Therefore, it is essential to develop tools to detect privacy leakage in Contrastive Language-Audio Pretr…
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
Shuirong Cao, Ruoxi Cheng, Zhiqiang Wang
LLMs can exhibit age biases, resulting in unequal treatment of individuals across age groups. While much research has addressed racial and gender biases, age bias remains little ex…