1 citations · 1 across the 1 of their papers we have counts for
1 paper
Xing Wu, Chaochen Gao, Meng Lin +3
Before entering the neural network, a token is generally converted to the corresponding one-hot representation, which is a discrete distribution of the vocabulary. Smoothed represe…