1 paper · 1 filter
Ha Young Kim, Niranjan Balasubramanian, Byungkon Kang
It has become common practice now to use random initialization schemes, rather than the pre-trained embeddings, when training transformer based models from scratch. Indeed, we find…