438 citations · 670 across the 14 of their papers we have counts for
4 papers · 1 filter
International Scientific Report on the Safety of Advanced AI (Interim Report)
Yoshua Bengio, Sören Mindermann, Daniel Privitera +41
This is the interim publication of the first International Scientific Report on the Safety of Advanced AI. The report synthesises the scientific understanding of general-purpose AI…
SoK: Watermarking for AI-Generated Content
Xuandong Zhao, Sam Gunn, Miranda Christ +11
As the outputs of generative AI (GenAI) techniques improve in quality, it becomes increasingly challenging to distinguish them from human-created content. Watermarking schemes are…
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
Jie Zhang, Christian Schlarmann, Kristina Nikolić +4
Ensemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermed…
Persistent Pre-Training Poisoning of LLMs
Yiming Zhang, Javier Rando, Ivan Evtimov +5
Large language models are pre-trained on uncurated text datasets consisting of trillions of tokens scraped from the Web. Prior work has shown that: (1) web-scraped pre-training dat…