1 paper · 1 filter
Karl Neergaard, Le Qiu, Emmanuele Chersoni
Single-prompt evaluations dominate current LLM benchmarking, yet they fail to capture the conversational dynamics where real-world harm occurs. In this study, we examined whether c…