1 paper · 1 filter
Jing-Jing Li, Valentina Pyatkin, Max Kleiman-Weiner +7
The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect…