1 paper
Ivan Dubrovsky, Anastasia Orlova, Illarion Iov +3
Benchmarking outcomes increasingly govern trust, selection, and deployment of LLMs, yet these evaluations remain vulnerable to semantically equivalent adversarial perturbations. Pr…