6 papers
You Only Charge Once 2.0 : A End-to-End Analog Computing-in-Memory Architecture with Reconfigurable Switched Capacitors
Zihao Xuan, Yewen Li, Jia Chen +3
Analog Computing-in-Memory (ACiM) accelerates deep neural networks by keeping weights inside memory arrays and executing dot products in the analog domain. However, modern ACiM acc…
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
Zihao Xuan, Jia Chen, Yewen Li +4
In this paper, we propose FusionCIM, an operator-fusion-driven compute-in-memory (CIM) accelerator architecture for efficient and scalable LLM inference, with three key innovations…
Towards Secure and Efficient DNN Accelerators via Hardware-Software Co-Design
Wei Xuan, Zihao Xuan, Rongliang Fu +8
The rapid deployment of deep neural network (DNN) accelerators in safety-critical domains such as autonomous vehicles, healthcare systems, and financial infrastructure necessitates…
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
Chao Han, Yijuan Liang, Zihao Xuan +3
The deployment of large language models (LLMs) in real-world applications is increasingly limited by their high inference cost. While recent advances in dynamic token-level computa…
SeDA: Secure and Efficient DNN Accelerators with Hardware/Software Synergy
Wei Xuan, Zhongrui Wang, Lang Feng +6
Ensuring the confidentiality and integrity of DNN accelerators is paramount across various scenarios spanning autonomous driving, healthcare, and finance. However, current security…
YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AI
Zihao Xuan, Yuxuan Yang, Wei Xuan +3
In this paper, we further explore the potential of analog in-memory computing (AiMC) and introduce an innovative artificial intelligence (AI) accelerator architecture named YOCO, f…