From the 1 of 1 linked paper with an AI index.
1 paper
Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal +1
The paper investigates how to suppress specific internal activations in large language models by optimizing only the input prompt, aiming to hide evaluation-awareness signals witho…