1 paper · 1 filter
Masha Fedzechkina, Eleonora Gualdoni, Rita Ramos +1
Recent work has shown that probing model internals can reveal a wealth of information not apparent from the model generations. This poses the risk of unintentional or malicious inf…