1 paper
Markus N. Rabe, Judith Clymo, Zheren Dong
We introduce a simple modification to the embedding layer. The key change is to infuse token embeddings with information about their spelling. Models trained with these embeddings…