2 papers
cs.LG2026
An Interpretable Latency Model for Speculative Decoding in LLM Serving
Linghao Kong, Megan Flynn, Michael Peng +3
Speculative decoding (SD) accelerates large language model (LLM) inference by using a smaller draft model to propose multiple tokens that are verified by a larger target model in p…
cs.CL2026
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
Maximillian Chen, Xuanming Zhang, Michael Peng +3
The rise of Internet of Things (IoT) devices in the physical world necessitates voice-based interfaces capable of handling complex user experiences. While modern Large Language Mod…