243 citations · 535 across the 8 of their papers we have counts for
5 papers · 1 filter
Watermarking Needs Input Repetition Masking
David Khachaturov, Robert Mullins, Ilia Shumailov +1
Recent advancements in Large Language Models (LLMs) raised concerns over potential misuse, such as for spreading misinformation. In response two counter measures emerged: machine l…
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz +31
We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…
Enabling certification of verification-agnostic networks via memory-efficient semidefinite programming
Sumanth Dathathri, Krishnamurthy Dvijotham, Alexey Kurakin +8
Convex relaxations have emerged as a promising approach for verifying desirable properties of neural networks like robustness to adversarial perturbations. Widely used Linear Progr…
Robust Constrained Reinforcement Learning for Continuous Control with Model Misspecification
Daniel J. Mankowitz, Dan A. Calian, Rae Jeong +5
Many real-world physical control systems are required to satisfy constraints upon deployment. Furthermore, real-world systems are often subject to effects such as non-stationarity,…
Detecting Adversarial Examples via Neural Fingerprinting
Sumanth Dathathri, Stephan Zheng, Tianwei Yin +2
Deep neural networks are vulnerable to adversarial examples, which dramatically alter model output using small input changes. We propose Neural Fingerprinting, a simple, yet effect…