From the 1 of 11 linked papers with an AI index.
1 paper · 1 filter
Usha Bhalla, Alex Oesterling, Claudio Mayrink Verdun +2
Translating the internal representations and computations of models into concepts that humans can understand is a key goal of interpretability. While recent dictionary learning met…