1 paper
Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail +7
Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult…