1 paper
Sullam Jeoung, Yubin Ge, Jana Diesner
Large Language Models (LLMs) have been observed to encode and perpetuate harmful associations present in the training data. We propose a theoretically grounded framework called Ste…