1 paper · 1 filter
Rachel Longjohn, Shang Wu, Saatvik Kher +2
It is increasingly important to evaluate how text generation systems based on large language models (LLMs) behave, such as their tendency to produce harmful output or their sensiti…