activity
20242026
collaborators

5 papers

cs.AR2026

FlexPosit: Tunable Fractional Precision for LLM Inference Accelerators

Yimin Gao, Liangtao Dai, Jun Yin +2

Large language models (LLMs) offer remarkable capabilities but impose prohibitive compute and energy costs. Quantization governs the trade-offs between accuracy and hardware effici…

cs.LG2025

CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs

Zhiyuan Ning, Jiawei Shao, Ruge Xu +4

Speculative decoding has become a widely adopted as an effective technique for lossless inference acceleration when deploying large language models (LLMs). While on-the-fly self-sp…

cs.AR2025

DAS-MP: Enabling High-Quality Macro Placement with Enhanced Dataflow Awareness

Xiaotian Zhao, Zixuan Li, Yichen Cai +3

Dataflow is a critical yet underexplored factor in automatic macro placement, which is becoming increasingly important for developing intelligent design automation techniques that…

cs.AR2025

WISP: Image Segmentation-Based Whitespace Diagnosis for Optimal Rectilinear Floorplanning

Xiaotian Zhao, Zixuan Li, Yichen Cai +1

The increasing number of rectilinear floorplans in modern chip designs presents significant challenges for traditional macro placers due to the additional complexity introduced by…

cs.LG2024

FAMES: Fast Approximate Multiplier Substitution for Mixed-Precision Quantized DNNs--Down to 2 Bits!

Yi Ren, Ruge Xu, Xinfei Guo +1

A widely-used technique in designing energy-efficient deep neural network (DNN) accelerators is quantization. Recent progress in this direction has reduced the bitwidths used in DN…