4 papers
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
Zhiyuan Ning, Jiawei Shao, Ruge Xu +4
Speculative decoding has become a widely adopted as an effective technique for lossless inference acceleration when deploying large language models (LLMs). While on-the-fly self-sp…
FAMES: Fast Approximate Multiplier Substitution for Mixed-Precision Quantized DNNs--Down to 2 Bits!
Yi Ren, Ruge Xu, Xinfei Guo +1
A widely-used technique in designing energy-efficient deep neural network (DNN) accelerators is quantization. Recent progress in this direction has reduced the bitwidths used in DN…
DAS-MP: Enabling High-Quality Macro Placement with Enhanced Dataflow Awareness
Xiaotian Zhao, Zixuan Li, Yichen Cai +3
Dataflow is a critical yet underexplored factor in automatic macro placement, which is becoming increasingly important for developing intelligent design automation techniques that…
WISP: Image Segmentation-Based Whitespace Diagnosis for Optimal Rectilinear Floorplanning
Xiaotian Zhao, Zixuan Li, Yichen Cai +1
The increasing number of rectilinear floorplans in modern chip designs presents significant challenges for traditional macro placers due to the additional complexity introduced by…