67 citations · 67 across the 2 of their papers we have counts for
2 papers
cs.CV2025
Astrea: A MOE-based Visual Understanding Model with Progressive Alignment
Xiaoda Yang, JunYu Lu, Hongshun Qiu +12
Vision-Language Models (VLMs) based on Mixture-of-Experts (MoE) architectures have emerged as a pivotal paradigm in multimodal understanding, offering a powerful framework for inte…
cs.LG2022★ 67 cited
Merak: An Efficient Distributed DNN Training Framework with Automated 3D Parallelism for Giant Foundation Models
Zhiquan Lai, Shengwei Li, Xudong Tang +5
Foundation models are becoming the dominant deep learning technologies. Pretraining a foundation model is always time-consumed due to the large scale of both the model parameter an…