3 papers
cs.AR2026
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
Songchen Ma, Hongyi Li, Weihao Zhang +8
Mixture-of-Experts is a promising approach for edge AI with low-batch inference. Yet, on-device deployments often face limited on-chip memory and severe workload imbalance; the pre…
cs.AR2025
CompAir: Synergizing Complementary PIMs and In-Transit NoC Computation for Efficient LLM Acceleration
Hongyi Li, Songchen Ma, Huanyu Qu +5
The rapid advancement of Large Language Models (LLMs) has revolutionized various aspects of human life, yet their immense computational and energy demands pose significant challeng…
cs.AR2025
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
Huanyu Qu, Weihao Zhang, Junfeng Lin +4
To efficiently support large-scale NNs, multi-level hardware, leveraging advanced integration and interconnection technologies, has emerged as a promising solution to counter the s…