1 paper · 1 filter
Xuran Li, Jingyi Wang
Large language models (LLMs) are increasingly deployed in real-world systems, yet they can produce toxic or biased outputs that undermine safety and trust. Post-hoc model repair pr…