Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Resurrecting saturated LLM benchmarks with adversarial encoding
Igor Ivanov, Dmitrii Volkov
Recent work showed that small changes in benchmark questions can reduce LLMs' reasoning and recall. We explore two such changes: pairing questions and adding more answer options, o…
cs.LG2024
Badllama 3: removing safety finetuning from Llama 3 in minutes
Dmitrii Volkov
We show that extensive LLM safety fine-tuning is easily subverted when an attacker has access to model weights. We evaluate three state-of-the-art fine-tuning methods-QLoRA, ReFT,…