2 papers
cs.DC2026
Fine-Tuning GPT-5 for GPU Kernel Generation
Ali Tehrani, Yahya Emara, Essam Wissam +4
Developing efficient GPU kernels is essential for scaling modern AI systems, yet it remains a complex task due to intricate hardware architectures and the need for specialized opti…
cs.CL2024
MobileQuant: Mobile-friendly Quantization for On-device Language Models
Fuwen Tan, Royson Lee, Åukasz Dudziak +5
Large language models (LLMs) have revolutionized language processing, delivering outstanding results across multiple applications. However, deploying LLMs on edge devices poses sev…