2 papers
cs.CL2025
Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
Jillian Xu, Dylan Zhou, Vinay Shukla +6
Chain-of-Thought (CoT) prompting often improves classification accuracy, but it introduces a significant throughput penalty with rationale generation (Wei et al., 2022; Cheng and V…
cs.CL2025
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
Zikun Li, Zhuofu Chen, Remi Delacourt +11
Modern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed const…