1 paper
Toufiq Parag, Ahmed Elgammal
The long-context capability of recent large transformer models can be surmised to rely on techniques such as attention/model parallelism, as well as hardware-level optimizations. W…