4 citations · 8 across the 5 of their papers we have counts for
1 paper · 2 filters
Shuai Zhang, Peng Zhang, Xindian Ma +3
Transformer has been widely-used in many Natural Language Processing (NLP) tasks and the scaled dot-product attention between tokens is a core module of Transformer. This attention…