243 citations · 439 across the 6 of their papers we have counts for
18 papers
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz +31
We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…
Challenges in Detoxifying Language Models
Johannes Welbl, Amelia Glaese, Jonathan Uesato +7
Large language models (LM) generate remarkably fluent text and can be efficiently adapted across NLP tasks. Measuring and guaranteeing the quality of generated text in terms of saf…
Towards Robust Image Classification Using Sequential Attention Models
Daniel Zoran, Mike Chrzanowski, Po-Sen Huang +3
In this paper we propose to augment a modern neural-network architecture with an attention model inspired by human perception. Specifically, we adversarially train and analyze a ne…
Achieving Robustness in the Wild via Adversarial Mixing with Disentangled Representations
Sven Gowal, Chongli Qin, Po-Sen Huang +4
Recent research has made the surprising finding that state-of-the-art deep learning models sometimes fail to generalize to small variations of the input. Adversarial training has b…
Reducing Sentiment Bias in Language Models via Counterfactual Evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang +6
Advances in language modeling architectures and the availability of large text corpora have driven progress in automatic text generation. While this results in models capable of ge…