Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
Weilin Wan, Jingtao Han, Weizhong Zhang +1
Scaling laws for Large Language Models govern macroscopic resource allocation, yet translating them into precise Mixture-of-Experts (MoE) architectural configurations remains an op…
cs.LG2025
Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
Weilin Wan, Fan Yi, Weizhong Zhang +2
Modern deep neural networks rely heavily on massive model weights and training samples, incurring substantial computational costs. Weight pruning and coreset selection are two emer…
cs.LG2025
Computational Budget Should Be Considered in Data Selection
Weilin Wan, Weizhong Zhang, Cheng Jin
Data selection improves computational efficiency by choosing informative subsets of training samples. However, existing methods ignore the compute budget, treating data selection a…