26 citations · 26 across the 1 of their papers we have counts for
1 paper · 1 filter
Zijie J. Wang, Robert Turko, Duen Horng Chau
Why do large pre-trained transformer-based models perform so well across a wide variety of NLP tasks? Recent research suggests the key may lie in multi-headed attention mechanism's…