1 paper
Ankit Kashyap
We present a Transformer architecture for long-context language modeling that combines global attention with two biologically inspired components: chunked local attention and a gat…