2 citations · 2 across the 3 of their papers we have counts for
3 papers
The Evaluation Game: Beyond Static LLM Benchmarking
Paul Wang, Jade Garcia-Bourrée, Anne-Marie Kermarrec +1
As jailbreaks, adversarially crafted inputs that bypass safety constraints, continue to be discovered in Large Language Models, practitioners increasingly rely on fine-tuning as a…
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
Diego Dorn, Alexandre Variengien, Charbel-Raphaël Segerie +1
Input-output safeguards are used to detect anomalies in the traces produced by Large Language Models (LLMs) systems. These detectors are at the core of diverse safety-critical appl…
Adversarial Imitation Learning On Aggregated Data
Pierre Le Pelletier de Woillemont, Rémi Labory, Vincent Corruble
Inverse Reinforcement Learning (IRL) learns an optimal policy, given some expert demonstrations, thus avoiding the need for the tedious process of specifying a suitable reward func…