Showing cs.ARShow all
3 papers · 1 filter
cs.AR2026
HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory Processing
Haochen Huang, Shuzhang Zhong, Shengxuan Qiu +8
Mixture-of-Experts (MoE) architectures have become a key technique for scaling Large Language Models (LLMs), enabling high model capacity with reduced computational cost. However,…
cs.AR2026
DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference
Jinqi Wen, Tong Xie, Runsheng Wang +1
Diffusion model deployment has been suffering from high energy consumption and inference latency despite its superior performance in visual generation tasks. Dynamic voltage and fr…
cs.AR2026
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
Zizhuo Fu, Yifan Zhou, Zhaoxin Lu +4
Multi-FPGA systems (MFS) are widely adopted for VLSI emulation and rapid prototyping. In an MFS, FPGAs connect only to a limited number of neighbors through bandwidth-constrained l…