most citedCatastrophic Jailbreak of Open-source LLMs via Exploiting Generation

11 citations · 31 across the 9 of their papers we have counts for

collaborators

8 papers

cs.CV20241 cited

ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty

Xindi Wu, Dingli Yu, Yangsibo Huang +2

Compositionality is a critical capability in Text-to-Image (T2I) models, as it reflects their ability to understand and combine multiple concepts from text descriptions. Existing e…

cs.CL20243 cited

MUSE: Machine Unlearning Six-Way Evaluation for Language Models

Weijia Shi, Jaechan Lee, Yangsibo Huang +7

Language models (LMs) are trained on vast amounts of text data, which may include private and copyrighted content. Data owners may request the removal of their data from a trained…

cs.AI202411 cited

A Safe Harbor for AI Evaluation and Red Teaming

Shayne Longpre, Sayash Kapoor, Kevin Klyman +20

Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems. However, the terms of service and enforcement strategies used by promi…

cs.LG2023

Sparsity-Preserving Differentially Private Training of Large Embedding Models

Badih Ghazi, Yangsibo Huang, Pritish Kamath +4

As the use of large embedding models in recommendation systems and language applications increases, concerns over user data privacy have also risen. DP-SGD, a training algorithm th…

cs.CL202311 cited

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Yangsibo Huang, Samyak Gupta, Mengzhou Xia +2

The rapid progress in open-source large language models (LLMs) is significantly advancing AI development. Extensive efforts have been made before model release to align their behav…

cs.LG2023

Learning across Data Owners with Joint Differential Privacy

Yangsibo Huang, Haotian Jiang, Daogao Liu +3

In this paper, we study the setting in which data owners train machine learning models collaboratively under a privacy notion called joint differential privacy [Kearns et al., 2018…