1 paper
Mokshit Surana, Archit Rathod, Akshaj Satishkumar
Large Language Models (LLMs) trained on web-scale corpora inherently absorb toxic patterns from their training data. This leads to toxic degeneration where even innocuous prompts c…