1 paper
Ahmed Abdelmuniem Abdalla Mohammed
Standard transformer architectures apply the same number of layers to every token regardless of contextual difficulty. We present Token-Selective Attention (TSA), a learned per-tok…