Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
Jiaming Yan, Jianchun Liu, Hongli Xu +1
Mixture-of-Experts (MoE) has emerged as a promising architecture for modern large language models (LLMs). However, massive parameters impose heavy GPU memory (i.e., VRAM) demands,…
cs.LG2024
Fair Differentiable Neural Network Architecture Search for Long-Tailed Data with Self-Supervised Learning
Jiaming Yan
Recent advancements in artificial intelligence (AI) have positioned deep learning (DL) as a pivotal technology in fields like computer vision, data mining, and natural language pro…