1 paper · 1 filter
Eduard Kapelko
Safety and controllability are critical for large language models. A central question is whether undesirable behaviors like deception are localized functions that can be removed, o…