1 paper
Namgyu Ho, Sangmin Bae, Taehyeon Kim +6
We introduce the Block Transformer which adopts hierarchical global-to-local modeling to autoregressive transformers to mitigate the inference bottlenecks associated with self-atte…