1 paper · 1 filter
Mokshit Surana, Archit Rathod, Akshaj Satishkumar
Large Language Models (LLMs) trained on web-scale corpora inherently absorb toxic patterns from their training data. This leads to toxic degeneration where even innocuous prompts c…