2 papers
cs.CV2026
FAR: Function-preserving Attention Replacement for IMC-friendly Inference
Yuxin Ren, Maxwell D Collins, Miao Hu +1
While transformers dominate modern vision and language models, their attention mechanism remains poorly suited for in-memory computing (IMC) devices due to intensive activation-to-…
cs.DB2025
FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference
Dongwei Wang, Zijie Liu, Song Wang +5
The Key-Value (KV) cache reading latency increases significantly with context lengths, hindering the efficiency of long-context LLM inference. To address this, previous works propo…