3 papers
cs.AR2026
Harvesting AI Computation at the Edge via Generic Approximation
Yihan Wang, Huiru Yan, Luxin Zhang +6
With the widespread adoption of AI in various IoT scenarios such as smart sensing and processing, AI chips have become a common component at the edge. These chips are typically spe…
cs.AR2024
COMET: Towards Partical W4A4KV4 LLMs Serving
Lian Liu, Haimeng Ren, Long Cheng +6
Quantization is a widely-used compression technology to reduce the overhead of serving large language models (LLMs) on terminal devices and in cloud data centers. However, prevalen…
cs.AR2024
CIM-MLC: A Multi-level Compilation Stack for Computing-In-Memory Accelerators
Songyun Qu, Shixin Zhao, Bing Li +4
In recent years, various computing-in-memory (CIM) processors have been presented, showing superior performance over traditional architectures. To unleash the potential of various…