1 paper
Jie Zhang, Meng Ding, Yang Liu +2
We present a novel approach for attacking black-box large language models (LLMs) by exploiting their ability to express confidence in natural language. Existing black-box attacks r…