2 papers
cs.CL2025
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
Jianghui Wang, Vinay Joshi, Saptarshi Majumder +7
The demand for AI-generated GPU kernels is rapidly growing, influenced by the need for scalable, hardware-optimized solutions in both industry and academia. As deep learning worklo…
cs.CL2025
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
Vinay Joshi, Pratik Prabhanjan Brahma, Zicheng Liu +1
The key-value (KV) cache in transformer models is a critical component for efficient decoding or inference, yet its memory demands scale poorly with sequence length, posing a major…