collaborators

6 papers

cs.AR2025

Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models

Chiyue Wei, Cong Guo, Junyao Zhang +8

Vision-Language Models (VLMs) have demonstrated strong performance on tasks such as video captioning and visual question answering. However, their growing scale and video-level inp…

cs.AR2025

Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication

Haoxuan Shan, Cong Guo, Chiyue Wei +4

The rapid scaling of large language models demands more efficient hardware. Quantization offers a promising trade-off between efficiency and performance. With ultra-low-bit quantiz…

cs.AR2025

AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems

Feng Cheng, Tunhou Zhang, Junyao Zhang +6

The performance bottleneck of deep-learning-based recommender systems resides in their backbone Deep Neural Networks. By integrating Processing-In-Memory~(PIM) architectures, resea…

cs.AR2025

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression

Feng Cheng, Cong Guo, Chiyue Wei +7

Large language models (LLMs) have demonstrated transformative capabilities across diverse artificial intelligence applications, yet their deployment is hindered by substantial memo…

cs.LG2024

Improving Routability Prediction via NAS Using a Smooth One-shot Augmented Predictor

Arjun Sridhar, Chen-Chia Chang, Junyao Zhang +1

Routability optimization in modern EDA tools has benefited greatly from using machine learning (ML) models. Constructing and optimizing the performance of ML models continues to be…

quant-ph2024

qGDP: Quantum Legalization and Detailed Placement for Superconducting Quantum Computers

Junyao Zhang, Guanglei Zhou, Feng Cheng +6

Noisy Intermediate-Scale Quantum (NISQ) computers are currently limited by their qubit numbers, which hampers progress towards fault-tolerant quantum computing. A major challenge i…