133 citations · 184 across the 17 of their papers we have counts for
4 papers · 1 filter
RePO: Bridging On-Policy Learning and Off-Policy Knowledge through Rephrasing Policy Optimization
Linxuan Xia, Xiaolong Yang, Yongyuan Chen +4
Aligning large language models (LLMs) on domain-specific data remains a fundamental challenge. Supervised fine-tuning (SFT) offers a straightforward way to inject domain knowledge…
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz +31
We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…
Attacking Adversarial Attacks as A Defense
Boxi Wu, Heng Pan, Li Shen +6
It is well known that adversarial attacks can fool deep neural networks with imperceptible perturbations. Although adversarial training significantly improves model robustness, fai…
Do Wider Neural Networks Really Help Adversarial Robustness?
Boxi Wu, Jinghui Chen, Deng Cai +2
Adversarial training is a powerful type of defense against adversarial examples. Previous empirical results suggest that adversarial training requires wider networks for better per…