11 papers
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
Xin Ye, Daning Cheng, Boyang Zhang +1
Training large-scale Mixture-of-Experts (MoE) models typically requires high-memory, high-bandwidth GPUs (e.g., A100), and their high cost has become a major barrier to large-model…
Rethinking Parameter Sharing as Graph Coloring for Structured Compression
Boyang Zhang, Daning Cheng, Yunquan Zhang
Modern deep models have massive parameter sizes, leading to high inference-time memory usage that limits practical deployment. Parameter sharing, a form of structured compression,…
MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
Jinhao Zhang, Yunquan Zhang, Boyang Zhang +2
Quantization method plays a crucial role in improving model efficiency and reducing deployment costs, enabling the widespread application of deep learning models on resource-constr…
A Unified Data-Driven Framework for Efficient Scientific Discovery
Tingxiong Xiao, Xinxin Song, Ziqian Wang +2
Scientific discovery drives progress across disciplines, from fundamental physics to industrial applications. However, identifying physical laws automatically from gathered dataset…
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
Zeyu Liu, Yan Li, Yunquan Zhang +6
Training large language models typically demands extensive GPU memory and substantial financial investment, which poses a barrier for many small- to medium-sized teams. In this pap…
SP2RINT: Spatially-Decoupled Physics-Inspired Progressive Inverse Optimization for Scalable, PDE-Constrained Meta-Optical Neural Network Training
Pingchuan Ma, Ziang Yin, Qi Jing +8
DONNs leverage light propagation for efficient analog AI and signal processing. Advances in nanophotonic fabrication and metasurface-based wavefront engineering have opened new pat…