3 papers
cs.AR2026
AUTOGATE: Automated Clock Gating via Toggling-Aware LLM-based RTL Rewriting
Yiting Wang, Chenhui Deng, Chia-Tung Ho +6
Fine-grain clock gating (FGCG) is among the most effective techniques for reducing dynamic power, yet current FGCG optimization flows remain largely manual. Recent LLM-based RTL op…
cs.AR2026
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
Jiayi Wang, Maohua Nie, Sin-Chen Lin +2
Commercial FPGAs, such as AMD Versal devices, increasingly incorporate AI engines that exploit low-precision packed-SIMD fused multiply-accumulate (FMA) to achieve proportional thr…
cs.AR2023
Duet: Creating Harmony between Processors and Embedded FPGAs
Ang Li, August Ning, David Wentzlaff
The demise of Moore's Law has led to the rise of hardware acceleration. However, the focus on accelerating stable algorithms in their entirety neglects the abundant fine-grained ac…