3 papers
cs.AR2026
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
Bowen Duan, Cong Guo, Chiyue Wei +8
Large Language Models (LLMs) have achieved impressive performance across diverse domains but remain inefficient during the autoregressive decoding phase. Unlike the prefill stage,…
cs.SE2026
Skilled AI Agents for Embedded and IoT Systems Development
Yiming Li, Yuhan Cheng, Mingchen Ma +6
Large language models (LLMs) and agentic systems have shown promise for automated software development, but applying them to hardware-in-the-loop (HIL) embedded and Internet-of-Thi…
cs.AR2025
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
Chiyue Wei, Cong Guo, Feng Cheng +4
Spiking Neural Networks (SNNs) are highly efficient due to their spike-based activation, which inherently produces bit-sparse computation patterns. Existing hardware implementation…