Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
Qi Fan, An Zou, Yehan Ma
Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a sub…
cs.CL2025
TimeBill: Time-Budgeted Inference for Large Language Models
Qi Fan, An Zou, Yehan Ma
Large Language Models (LLMs) are increasingly deployed in time-critical systems, such as robotics, autonomous driving, embodied intelligence, and industrial automation, where gener…