3 papers
cs.LG2025
Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation
Zhenyu Liu, Yunzhen Liu, Zehao Fan +5
Mixture-of-Experts (MoE) models scale capacity via sparse activation but stress memory and bandwidth. Offloading alleviates GPU memory by fetching experts on demand, yet token-leve…
cs.LG2025
Ralts: Robust Aggregation for Enhancing Graph Neural Network Resilience on Bit-flip Errors
Wencheng Zou, Nan Wu
Graph neural networks (GNNs) have been widely applied in safety-critical applications, such as financial and medical networks, in which compromised predictions may cause catastroph…
cs.LG2024
A Benchmark on Directed Graph Representation Learning in Hardware Designs
Haoyu Wang, Yinan Huang, Nan Wu +1
To keep pace with the rapid advancements in design complexity within modern computing systems, directed graph representation learning (DGRL) has become crucial, particularly for en…