1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.DC2025
AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
Huawei Bai, Yifan Huang, Wenqi Shi +4
The training efficiency and scalability of language models on massive clusters currently remain a critical bottleneck. Mainstream approaches like ND parallelism are often cumbersom…
cs.CL2025★ 1 cited
Seed-Coder: Let the Code Model Curate Data for Itself
ByteDance Seed, Yuyu Zhang, Jing Su +24
Code data in large language model (LLM) pretraining is recognized crucial not only for code-related tasks but also for enhancing general intelligence of LLMs. Current open-source L…
cs.SE2025
A Vulnerability Code Intent Summary Dataset
Yifan Huang, Weisong Sun, Yubin Qu
In the era of Large Language Models (LLMs), the code summarization technique boosts a lot, along with the emergence of many new significant works. However, the potential of code su…