From the 1 of 30 linked papers with an AI index.
30 papers
TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models
Alireza Dehghanpour Farashah, Zhuan Shi, Negar Rostamzadeh +1
Text-to-image diffusion models can be misused to generate harmful content through adversarial or paraphrased prompts that bypass built-in safety mechanisms. Existing concept erasur…
IP Protection in the Era of Visual Generative AI: A Survey
Zhuan Shi, Shunchang Liu, Alireza Dehghanpour Farashah +8
The rapid evolution of visual generative AI has introduced a wide range of intellectual property risks, spanning the unauthorized learning, reproduction, extraction, misuse, and re…
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
Marylou Fauchard, Florian Carichon, Margarida Carvalho +1
The paper studies how misaligned objectives in large‑language‑model powered multi‑agent systems affect performance in the social deduction game Werewolf, showing that hidden object…
Revisiting the Neural Tangent Kernel: the role of large width and depth
William St-Arnaud, Margarida Carvalho, Golnoosh Farnadi
Overparameterized fully-connected neural networks have been shown to behave like kernel models when trained with gradient descent, assuming standard scaling conditions on the width…
Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior
Ali khalil, Aly M. Kassem, Mohamed Abdelrazek +3
We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using…
IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts
Ayana Hussain, Soumya Sharma, Golnoosh Farnadi +3
Large language models (LLMs) are becoming widely deployed as personal AI assistants with access to sensitive user data, making privacy a major challenge for their design and evalua…