4 papers · 1 filter
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
Dhananjay Ashok, Ruth-Ann Armstrong, Jonathan May
Language Models (LMs) may acquire harmful knowledge, and yet feign ignorance of these topics when under audit. Inspired by the recent discovery of deception-related behaviour patte…
Language Models Can Predict Their Own Behavior
Dhananjay Ashok, Jonathan May
The text produced by language models (LMs) can exhibit specific `behaviors,' such as a failure to follow alignment training, that we hope to detect and react to during deployment.…
A Little Human Data Goes A Long Way
Dhananjay Ashok, Jonathan May
Faced with an expensive human annotation process, creators of NLP systems increasingly turn to synthetic data generation. While this method shows promise, the extent to which synth…
Controllable Text Generation in the Instruction-Tuning Era
Dhananjay Ashok, Barnabas Poczos
While most research on controllable text generation has focused on steering base Language Models, the emerging instruction-tuning and prompting paradigm offers an alternate approac…