3 citations · 4 across the 3 of their papers we have counts for
4 papers · 1 filter
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
Zeke Wang, Jie Zhang, Hongjing Huang +10
Modern data analytics requires a huge amount of computing power and processes a massive amount of data. At the same time, the underlying computing platform is becoming much more he…
DisDP: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel Training
Mo Sun, Zihan Yang, Changyue Liao +5
Model-sharded data parallelism (MSDP), e.g., ZeRO, evenly shards the model states across all GPUs, and thus has been widely adopted by LLM pre-training, such as Llama and DeepSeek,…
LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
Changyue Liao, Mo Sun, Zihan Yang +5
Nowadays, AI researchers become more and more interested in fine-tuning a pre-trained LLM, whose size has grown to up to over 100B parameters, for their downstream tasks. One appro…
Helios: An Efficient Out-of-core GNN Training System on Terabyte-scale Graphs with In-memory Performance
Jie Sun, Mo Sun, Zheng Zhang +6
Training graph neural networks (GNNs) on large-scale graph data holds immense promise for numerous real-world applications but remains a great challenge. Several disk-based GNN sys…