1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Liu Chen, Meysam Asgari
A deep Transformer model with good evaluation score does not mean each subnetwork (a.k.a transformer block) learns reasonable representation. Diagnosing abnormal representation and…