Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization
Gunjun Lee, Sehwan Son, Younjoo Lee +2
Weight-only post-training quantization (PTQ) enables the deployment of large language models under tight memory budgets, but accuracy often collapses at 2-3 bits. Existing backprop…
cs.LG2025
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
Gunjun Lee, Jiwon Kim, Jaiyoung Park +2
Large Language Model (LLM) inference in production must meet stringent service-level objectives for both time-to-first-token (TTFT) and time-between-token (TBT) while maximizing th…