2 papers
cs.LG2026
Deep Kernel Fusion for Transformers
Zixi Zhang, Zhiwen Mo, Yiren Zhao +1
Agentic LLM inference with long contexts is increasingly limited by memory bandwidth rather than compute. In this setting, SwiGLU MLP blocks, whose large weights exceed cache capac…
cs.LG2025
Hardware and Software Platform Inference
Cheng Zhang, Hanna Foerster, Robert D. Mullins +2
It is now a common business practice to buy access to large language model (LLM) inference rather than self-host, because of significant upfront hardware infrastructure and energy…