1 paper · 1 filter
Hiskias Dingeto
Natural-language autoencoders score explanations of hidden activations by reconstruction. An explanation is deemed faithful if the activation can be regenerated from it. The test i…