3 papers
cs.LG2026
KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning
Kris Shengjun Dong, Sahil Modi, Dima Nikiforov +4
Optimizing CUDA code across multiple generations of GPU architectures is challenging, as achieving peak performance requires an extensive exploration of an increasingly complex, ha…
cs.RO2024
Characterizing and Optimizing Real-Time Optimal Control for Embedded SoCs
Kris Shengjun Dong, Dima Nikiforov, Widyadewi Soedarmadji +4
Resource-limited robots face significant challenges in executing computationally intensive tasks, such as locomotion and manipulation, particularly for real-time optimal control al…
cs.AR2024
LLM-Aided Compilation for Tensor Accelerators
Charles Hong, Sahil Bhatia, Altan Haan +4
Hardware accelerators, in particular accelerators for tensor processing, have many potential application domains. However, they currently lack the software infrastructure to suppor…