2 papers
cs.AI2026
Rethinking Heterogeneous System Disaggregation for Subquadratic Attention
Arya Tschand, Yaosheng Fu, Vikram Sharma Mailthody +6
Frontier language models are more aggressively using subquadratic attention to reduce the memory footprint and compute requirements during inference while still delivering frontier…
cs.LG2026
SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
Edward Lin, Sahil Modi, Siva Kumar Sastry Hari +30
As agentic AI systems become increasingly capable of generating and optimizing GPU kernels, progress is constrained by benchmarks that reward speedup over software baselines rather…