4 papers
UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems
Mingming Ha, Guanchen Wang, Linxun Chen +9
In recent years, the scaling laws of recommendation models have attracted increasing attention, which govern the relationship between performance and parameters/FLOPs of recommende…
Scalable Analytic Classifiers with Associative Drift Compensation for Class-Incremental Learning of Vision Transformers
Xuan Rao, Mingming Ha, Bo Zhao +2
Class-incremental learning (CIL) with Vision Transformers (ViTs) faces a major computational bottleneck during the classifier reconstruction phase, where most existing methods rely…
Compensating Distribution Drifts in Class-incremental Learning of Pre-trained Vision Transformers
Xuan Rao, Simian Xu, Zheng Li +4
Recent advances have shown that sequential fine-tuning (SeqFT) of pre-trained vision transformers (ViTs), followed by classifier refinement using approximate distributions of class…
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
Xiaodong Chen, Mingming Ha, Zhenzhong Lan +2
The Mixture-of-Experts (MoE) architecture has become a predominant paradigm for scaling large language models (LLMs). Despite offering strong performance and computational efficien…