1 paper
Tian Chen, Mingheng Mi, Pu Wang +2
Large language models now evolve faster than production inference systems can be ported and optimized. New releases change attention, MoE routing, quantization formats, KV-cache la…