2 papers
cs.LG2026
Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving
Jae Gon Kim, Donghoon Yoo, Hanyul Ryu +3
Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference…
cs.LG2026
ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling
Sihwa Lee, Janghwan Lee, Donghoon Yoo +4
Large reasoning models (LRMs) improve complex problem-solving by generating long intermediate reasoning traces, but this substantially increases inference costs. NVFP4 inference of…