1 paper
Chris Cameron, Wangzheng Wang, Nikita Ivanov +3
Looped transformers scale computational depth without increasing parameter count by repeatedly applying a shared transformer block and can be used for iterative refinement, where e…