13 papers
MANCE: Manifold Aware Concept Erasure
Matan Avitan, Yoav Goldberg, Yanai Elazar
Concept erasure aims to remove a target concept from a representation while preserving the other information encoded in it. This is difficult because representations encode many co…
Benchmarking Agentic Review Systems
Dang Nguyen, Wanqing Hao, Yanai Elazar +1
A new class of agentic review systems are emerging as a remedy to the pressure placed on peer review systems by AI-assisted research, but it is unclear how they should be evaluated…
Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
Rahul Nadkarni, Yanai Elazar, Hila Gonen +1
We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for intervening on data batches -- i.e., ``…
Calibrating Large Language Models with Sample Consistency
Qing Lyu, Kumar Shridhar, Chaitanya Malaviya +6
Accurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and…
LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
Yanai Elazar, Maria Antoniak
ArXiv recently prohibited the upload of unpublished review papers to its servers in the Computer Science domain, citing a high prevalence of LLM-generated content in these categori…
How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models
Sahil Verma, Royi Rassin, Arnav Das +6
Text-to-image models are trained using large datasets of image-text pairs collected from the internet. These datasets often include copyrighted and private images. Training models…