most citedMoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

4 citations · 4 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI20244 cited

MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Xi Victoria Lin, Akshat Shrivastava, Liang Luo +5

We introduce MoMa, a novel modality-aware mixture-of-experts (MoE) architecture designed for pre-training mixed-modal, early-fusion language models. MoMa processes images and text…

physics.med-ph2023

X-ray computed tomography reconstruction algorithm for refractive index gradient

Keliang Liao, Qili He, Panyun Li +2

The aim of this research is to reconstruct the 3D X-ray refractive index gradient maps by the proposed vector Radon transform and its inverse, assuming that the small-angle deviati…

cs.DC2023

P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAs

Hongjing Huang, Yingtao Li, Jie Sun +5

Generalized linear models (GLMs) are a widely utilized family of machine learning models in real-world applications. As data size increases, it is essential to perform efficient di…

cs.LG2023

Pre-train and Search: Efficient Embedding Table Sharding with Pre-trained Neural Cost Models

Daochen Zha, Louis Feng, Liang Luo +8

Sharding a large machine learning model across multiple devices to balance the costs is important in distributed training. This is challenging because partitioning is NP-hard, and…

cs.LG2023

Self-discipline on multiple channels

Jiutian Zhao, Liang Luo, Hao Wang

Self-distillation relies on its own information to improve the generalization ability of the model and has a bright future. Existing self-distillation methods either require additi…