10 papers
Decoding the Skew: Distribution-Aware MoE Inference with Adaptive Kernel Dispatch
En-Ming Huang, An-Cheng Chang, Bai-Cheng Jeng +2
Mixture-of-Experts (MoE) inference consists of sparse expert GEMMs whose shapes vary with the runtime routing distribution. Existing serving systems typically select fused-MoE kern…
Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation
En-Ming Huang, Yu-Hung Kao, Ren-Hao Deng +10
Automated testbench generation has become a critical bottleneck in large language model (LLM)-driven Register Transfer Level (RTL) workflows, where large numbers of candidate desig…
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation
Mu-Chi Chen, Po-Hsuan Huang, Yu-Hung Kao +6
Recent advances in large language models have improved code generation, but their use in hardware description languages is still limited. Moreover, training data and testbenches fo…
Large-Scale Quantum Circuit Simulation on HPC Cluster via Cache Blocking, Boosting, and Gate Fusion Optimization
Chuan-Chi Wang, Yan-Jie Wang, Chia-Heng Tu +1
Quantum circuit simulation is crucial for the development of quantum algorithms, particularly given the high cost and noise limitations of physical quantum hardware. While full-sta…
Warp-STAR: High-performance, Differentiable GPU-Accelerated Static Timing Analysis through Warp-oriented Parallel Orchestration
En-Ming Huang, Shih-Hao Hung
Static timing analysis (STA) is crucial for Electronic Design Automation (EDA) flows but remains a computational bottleneck. While existing GPU-based STA engines are faster than CP…
ParaQAOA: Efficient Parallel Divide-and-Conquer QAOA for Large-Scale Max-Cut Problems Beyond 10,000 Vertices
Po-Hsuan Huang, Xie-Ru Li, Chi Chuang +2
Quantum Approximate Optimization Algorithm (QAOA) has emerged as a promising solution for combinatorial optimization problems using a hybrid quantum-classical framework. Among comb…