5 citations · 5 across the 1 of their papers we have counts for
1 paper · 1 filter
Jeongin Bae, Baeseong Park, Gunho Park +7
Transformer attention is typically implemented using softmax normalization, which enforces attention weights with unit sum normalization. While effective in many settings, this con…