2 papers
cs.MA2026
Automatic Model-Hardware Co-Adaptation for Heterogeneous AI Accelerators
Tian Chen, Mingheng Mi, Pu Wang +2
Large language models now evolve faster than production inference systems can be ported and optimized. New releases change attention, MoE routing, quantization formats, KV-cache la…
cs.CL2024
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
Ranajoy Sadhukhan, Jian Chen, Zhuoming Chen +7
Large Language Models (LLMs) have become more prevalent in long-context applications such as interactive chatbots, document analysis, and agent workflows, but it is challenging to…