Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026★ 4 cited
Tensor Product Attention Is All You Need
Yifan Zhang, Yifeng Liu, Huizhuo Yuan +4
Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during inference. In this pape…
cs.CL2025
Causal Attention with Lookahead Keys
Zhuoqing Song, Peng Sun, Huizhuo Yuan +1
In standard causal attention, each token's query, key, and value (QKV) are static and encode only preceding context. We introduce CAuSal aTtention with Lookahead kEys (CASTLE), an…