3 papers
cs.DC2025
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
Ruifan Chu, Anbang Wang, Xiuxiu Bai +2
In high-performance computing, hotspot GPU kernels are primary bottlenecks, and expert manual tuning is costly and hard to port. Large language model methods often assume kernels c…
cs.OS2025
Taiji: A DPU Memory Elasticity Solution for In-production Cloud Environments
Hao Zheng, Longxiang Wang, Yun Xu +22
The growth of cloud computing drives data centers toward higher density and efficiency. Data processing units (DPUs) enhance server network and storage performance but face challen…
cs.OS2025
Vmem: A Lightweight Hot-Upgradable Memory Management for In-production Cloud Environment
Hao Zheng, Qiang Wang, Longxiang Wang +16
Traditional memory management suffers from metadata overhead, architectural complexity, and stability degradation, problems intensified in cloud environments. Existing software/har…