7 papers
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
Haoxuan Shan, Cong Guo, Chiyue Wei +4
The rapid scaling of large language models demands more efficient hardware. Quantization offers a promising trade-off between efficiency and performance. With ultra-low-bit quantiz…
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
Feng Cheng, Tunhou Zhang, Junyao Zhang +6
The performance bottleneck of deep-learning-based recommender systems resides in their backbone Deep Neural Networks. By integrating Processing-In-Memory~(PIM) architectures, resea…
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
Feng Cheng, Cong Guo, Chiyue Wei +7
Large language models (LLMs) have demonstrated transformative capabilities across diverse artificial intelligence applications, yet their deployment is hindered by substantial memo…
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
Chiyue Wei, Cong Guo, Feng Cheng +4
Spiking Neural Networks (SNNs) are highly efficient due to their spike-based activation, which inherently produces bit-sparse computation patterns. Existing hardware implementation…
Towards Automated Model Design on Recommender Systems
Tunhou Zhang, Dehua Cheng, Yuchen He +10
The increasing popularity of deep learning models has created new opportunities for developing AI-based recommender systems. Designing recommender systems using deep neural network…
qGDP: Quantum Legalization and Detailed Placement for Superconducting Quantum Computers
Junyao Zhang, Guanglei Zhou, Feng Cheng +6
Noisy Intermediate-Scale Quantum (NISQ) computers are currently limited by their qubit numbers, which hampers progress towards fault-tolerant quantum computing. A major challenge i…