1 paper · 1 filter
Yawei Liu
Transformer-based pre-trained language models (PLMs) excel in text classification but suffer from attention dilution and attention sink effects, forcing models to over-focus on tas…