2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Zhiyuan Wang, Xuan Luo, Sirui Zeng +1
Decoder-only Transformers compute attention over the KV cache of preceding tokens. Keys (and Values) are typically represented with the same dimensionality, regardless of its dista…