2 papers
cs.AI2026
Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent
Lingyun Yang, Yuxiao Wang, Shenghao Liang +8
Existing GPU kernel generation benchmarks draw problems from synthetic or curated sources that diverge from deployed workloads. We present Atrex-Bench, a benchmark whose 30 operato…
cs.DC2025
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
Yuxiao Wang, Yuedong Xu, Qingyang Duan +4
The rapid growth of large language models (LLMs) and the continuous release of new GPU products have significantly increased the demand for distributed training across heterogeneou…