From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent
Lingyun Yang, Yuxiao Wang, Shenghao Liang +8
The paper introduces Atrex-Bench, a trace-driven GPU kernel benchmark derived from real production inference workloads, and evaluates LLM-generated kernels, revealing a large perfo…
cs.DC2025
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
Yuxiao Wang, Yuedong Xu, Qingyang Duan +4
The rapid growth of large language models (LLMs) and the continuous release of new GPU products have significantly increased the demand for distributed training across heterogeneou…