2 papers
cs.DC2026
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
Joon Ha Kim, Geon-Woo Kim, Anoop Rachakonda +1
Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, since no single choice performs…
cs.DC2025
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
Shaoyu Wang, Guangrong He, Geon-Woo Kim +2
Mixture-of-Experts (MoE) architectures offer the promise of larger model capacity without the prohibitive costs of fully dense designs. However, in real-world inference serving, lo…