4 citations · 4 across the 6 of their papers we have counts for
Showing 2025Show all
3 papers · 1 filter
cs.LG2025
Group Representational Position Encoding
Yifan Zhang, Zixiang Chen, Yifeng Liu +6
We present GRAPE (Group Representational Position Encoding), a unified framework for positional encoding based on group actions. GRAPE unifies two families of mechanisms: (i) multi…
cs.LG2025
Higher-order Linear Attention
Yifan Zhang, Zhen Qin, Mengdi Wang +1
The quadratic cost of scaled dot-product attention is a central obstacle to scaling autoregressive language models to long contexts. Linear-time attention and State Space Models (S…
cs.CL2025★ 4 cited
Tensor Product Attention Is All You Need
Yifan Zhang, Yifeng Liu, Huizhuo Yuan +4
Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during inference. In this pape…