2 papers
cs.DC2026
SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System
Shiyang Li, Guangyan Sun, Jinwei Tang +3
Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the input sparsity pattern and execu…
cs.LG2026
CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs
Shiyang Li, Zijian Zhang, Guangyan Sun +5
Large language models show promise for automated CUDA programming, however even the strongest coding models (e.g., Claude-Opus-4.6) may still fall short of expert-level, architectu…