2 citations · 2 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
The Evaluation Game: Beyond Static LLM Benchmarking
Paul Wang, Jade Garcia-Bourrée, Anne-Marie Kermarrec +1
As jailbreaks, adversarially crafted inputs that bypass safety constraints, continue to be discovered in Large Language Models, practitioners increasingly rely on fine-tuning as a…
cs.LG2023
Adversarial Imitation Learning On Aggregated Data
Pierre Le Pelletier de Woillemont, Rémi Labory, Vincent Corruble
Inverse Reinforcement Learning (IRL) learns an optimal policy, given some expert demonstrations, thus avoiding the need for the tedious process of specifying a suitable reward func…