Showing cs.ARShow all
3 papers · 1 filter
cs.AR2026
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
Bowen Duan, Cong Guo, Chiyue Wei +8
Large Language Models (LLMs) have achieved impressive performance across diverse domains but remain inefficient during the autoregressive decoding phase. Unlike the prefill stage,…
cs.AR2025
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
Chiyue Wei, Cong Guo, Feng Cheng +4
Spiking Neural Networks (SNNs) are highly efficient due to their spike-based activation, which inherently produces bit-sparse computation patterns. Existing hardware implementation…
cs.AR2023
NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data Processing
Yitu Wang, Shiyu Li, Qilin Zheng +5
Approximate nearest neighbor search (ANNS) is a key retrieval technique for vector database and many data center applications, such as person re-identification and recommendation s…