5 papers
MANCE: Manifold Aware Concept Erasure
Matan Avitan, Yoav Goldberg, Yanai Elazar
Concept erasure aims to remove a target concept from a representation while preserving the other information encoded in it. This is difficult because representations encode many co…
GRADE: Quantifying Sample Diversity in Text-to-Image Models
Royi Rassin, Aviv Slobodkin, Shauli Ravfogel +2
We introduce GRADE, an automatic method for quantifying sample diversity in text-to-image models. Our method leverages the world knowledge embedded in large language models and vis…
A Practical Method for Generating String Counterfactuals
Matan Avitan, Ryan Cotterell, Yoav Goldberg +1
Interventions targeting the representation space of language models (LMs) have emerged as an effective means to influence model behavior. Such methods are employed, for example, to…
Linear Adversarial Concept Erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg +1
Modern neural models trained on textual data rely on pre-trained representations that emerge without direct supervision. As these representations are increasingly being used in rea…
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
Amir DN Cohen, Shauli Ravfogel, Shaltiel Shmidman +1
In few-shot relation classification (FSRC), models must generalize to novel relations with only a few labeled examples. While much of the recent progress in NLP has focused on scal…