13 citations · 13 across the 3 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2022★ 13 cited
Characterizing Intrinsic Compositionality in Transformers with Tree Projections
Shikhar Murty, Pratyusha Sharma, Jacob Andreas +1
When trained on language data, do transformers learn some arbitrary computation that utilizes the full capacity of the architecture or do they learn a simpler, tree-like computatio…
cs.CL2021
How Do Neural Sequence Models Generalize? Local and Global Context Cues for Out-of-Distribution Prediction
Anthony Bau, Jacob Andreas
After a neural sequence model encounters an unexpected token, can its behavior be predicted? We show that RNN and transformer language models exhibit structured, consistent general…