◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Pengcheng Zhang

4 papers hereh-index 5287 citations10 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.DC4
same name
  • Pengcheng Zhang — 6 papers, h 8
  • Pengcheng Zhang — 6 papers, h 3
  • Pengcheng Zhang — 4 papers, h 2
  • Pengcheng Zhang — 3 papers
  • Pengcheng Zhang — 3 papers, h 6
  • Pengcheng Zhang — 2 papers, h 2

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20232025
most citedDxPU: Large Scale Disaggregated GPU Pools in the Datacenter

8 citations · 8 across the 1 of their papers we have counts for

collaborators

4 papers

cs.DC2025

EROICA: Online Performance Troubleshooting for Large-scale Model Training

Yu Guan, Zhiyu Yin, Haoyu Chen +11

Troubleshooting performance problems of large model training (LMT) is immensely challenging, due to unprecedented scales of modern GPU clusters, the complexity of software-hardware…

cs.DC2024

TrainMover: An Interruption-Resilient Runtime for ML Training

ChonLam Lao, Jiaqi Gao, Jiamin Cao +13

Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solutions like checkpoint-restart or runtime r…

cs.DC2024

Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization

Jianbo Dong, Bin Luo, Jun Zhang +22

The emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single mode…

cs.DC2023★ 8 cited

DxPU: Large Scale Disaggregated GPU Pools in the Datacenter

Bowen He, Xiao Zheng, Yuan Chen +9

The rapid adoption of AI and convenience offered by cloud services have resulted in the growing demands for GPUs in the cloud. Generally, GPUs are physically attached to host serve…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.