1 citations · 1 across the 1 of their papers we have counts for
1 paper
Allison Hegel, Marina Shah, Genevieve Peaslee +2
Large, pre-trained transformer models like BERT have achieved state-of-the-art results on document understanding tasks, but most implementations can only consider 512 tokens at a t…