2 papers
cs.PL2026
AgentCompile: An LLM-Guided Compiler for Direct CUDA Inference
Xuanzhe Li, Ziyan Weng, Zhiyu Zhu +1
Transformer inference increasingly relies on specialized compiler and runtime support, while recent LLMs can generate nontrivial CUDA kernels. However, unconstrained generation gua…
cs.AR2024
SigDLA: A Deep Learning Accelerator Extension for Signal Processing
Fangfa Fu, Wenyu Zhang, Zesong Jiang +8
Deep learning and signal processing are closely correlated in many IoT scenarios such as anomaly detection to empower intelligence of things. Many IoT processors utilize digital si…