3 papers
cs.CL2026
DRTriton: Large-Scale Synthetic Data Driven Reinforcement Learning for Triton Kernel Generation
Siqi Guo, Ming Lin, Tianbao Yang
Developing efficient CUDA kernels is a fundamental yet challenging task in the generative AI industry. Recent research leverages Large Language Models (LLMs) to automatically conve…
cs.AI2024
Search for Efficient Large Language Models
Xuan Shen, Pu Zhao, Yifan Gong +7
Large Language Models (LLMs) have long held sway in the realms of artificial intelligence research. Numerous efficient techniques, including weight pruning, quantization, and disti…
cs.LG2023
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
Xuan Shen, Peiyan Dong, Lei Lu +5
Large Language Models (LLMs) stand out for their impressive performance in intricate language modeling tasks. However, their demanding computational and memory needs pose obstacles…