From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory
Yizhe Chen, Wenshuai Yao, Saiya Wang +6
Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade low-bit quantized models. Existing CIM-or…
cs.CL2026
Recall Before You Rank: Similarity-Guided Top- Reuse for Efficient Long-Context Attention
Wenshuai Yao, Wenyong Zhou, Hanyong Shao +5
The paper proposes ReTopK, a training‑free technique that speeds up dynamic top‑K sparse attention for long‑context language models by reusing supports from historically similar qu…