1 paper · 1 filter
Ruining Li, Gabrijel Boduljak, Jensen +1
It is a widely known issue that Transformers, when trained on shorter sequences, fail to generalize robustly to longer ones at test time. This raises the question of whether Transf…