4 citations · 4 across the 5 of their papers we have counts for
5 papers
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
Xi Victoria Lin, Akshat Shrivastava, Liang Luo +5
We introduce MoMa, a novel modality-aware mixture-of-experts (MoE) architecture designed for pre-training mixed-modal, early-fusion language models. MoMa processes images and text…
X-ray computed tomography reconstruction algorithm for refractive index gradient
Keliang Liao, Qili He, Panyun Li +2
The aim of this research is to reconstruct the 3D X-ray refractive index gradient maps by the proposed vector Radon transform and its inverse, assuming that the small-angle deviati…
P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs
Hongjing Huang, Yingtao Li, Jie Sun +5
Generalized linear models (GLMs) are a widely utilized family of machine learning models in real-world applications. As data size increases, it is essential to perform efficient di…
Pre-train and Search: Efficient Embedding Table Sharding with Pre-trained Neural Cost Models
Daochen Zha, Louis Feng, Liang Luo +8
Sharding a large machine learning model across multiple devices to balance the costs is important in distributed training. This is challenging because partitioning is NP-hard, and…
Self-discipline on multiple channels
Jiutian Zhao, Liang Luo, Hao Wang
Self-distillation relies on its own information to improve the generalization ability of the model and has a bright future. Existing self-distillation methods either require additi…