8 papers
CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation
Yuanpeng Zhang, YuXuan Wu, Yitong Xiao +6
The paper introduces CODA, a hardware-software co-designed architecture that separates compute and cache operations for edge video diffusion models, using near‑memory processing to…
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
Cong Li, Chenhao Xue, Yi Ren +11
Large language models (LLMs) exhibit memory-intensive behavior during decoding, making it a key bottleneck in LLM inference. To accelerate decoding execution, hybrid-bonding-based…
CellE: Automated Standard Cell Library Extension via Equality Saturation
Yi Ren, Yukun Wang, Xiang Meng +6
Automated standard cell library extension is crucial for maximizing Quality of Results (QoR) in modern VLSI design. We introduce CellE, a novel framework that leverages formal meth…
Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
Yi Ren, Baokang Peng, Chenhao Xue +6
With the diminishing return from Moore's Law, system-technology co-optimization (STCO) has emerged as a promising approach to sustain the scaling trends in the VLSI industry. By br…
AC-Refiner: Efficient Arithmetic Circuit Optimization Using Conditional Diffusion Models
Chenhao Xue, Kezhi Li, Jiaxing Zhang +7
Arithmetic circuits, such as adders and multipliers, are fundamental components of digital systems, directly impacting the performance, power efficiency, and area footprint. Howeve…
FAMES: Fast Approximate Multiplier Substitution for Mixed-Precision Quantized DNNs--Down to 2 Bits!
Yi Ren, Ruge Xu, Xinfei Guo +1
A widely-used technique in designing energy-efficient deep neural network (DNN) accelerators is quantization. Recent progress in this direction has reduced the bitwidths used in DN…