26 citations · 43 across the 4 of their papers we have counts for
1 paper · 1 filter
Shikhar Murty, Pratyusha Sharma, Jacob Andreas +1
When trained on language data, do transformers learn some arbitrary computation that utilizes the full capacity of the architecture or do they learn a simpler, tree-like computatio…