3 papers
cs.AR2026
Decoding the Skew: Distribution-Aware MoE Inference with Adaptive Kernel Dispatch
En-Ming Huang, An-Cheng Chang, Bai-Cheng Jeng +2
Mixture-of-Experts (MoE) inference consists of sparse expert GEMMs whose shapes vary with the runtime routing distribution. Existing serving systems typically select fused-MoE kern…
cs.AR2026
SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation
Mu-Chi Chen, Yu-Hung Kao, Po-Hsuan Huang +10
Large language models (LLMs) have recently emerged as a promising approach for automating Verilog code generation; however, existing methods primarily emphasize syntactic correctne…
cs.LG2025
ElaLoRA: Elastic & Learnable Low-Rank Adaptation for Efficient Model Fine-Tuning
Huandong Chang, Zicheng Ma, Mingyuan Ma +4
Low-Rank Adaptation (LoRA) has become a widely adopted technique for fine-tuning large-scale pre-trained models with minimal parameter updates. However, existing methods rely on fi…