2 papers
cs.LG2025
From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
Junfeng Gong, Zhiyi Wei, Junying Chen +2
Despite significant evolution of CUDA programming and domain-specific libraries, effectively utilizing GPUs with massively parallel engines remains difficult. Large language models…
cs.AR2025
Large Processor Chip Model
Kaiyan Chang, Mingzhi Chen, Yunji Chen +40
Computer System Architecture serves as a crucial bridge between software applications and the underlying hardware, encompassing components like compilers, CPUs, coprocessors, and R…