Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
SOMA: Efficient Multi-turn LLM Serving via Small Language Model
Xueqi Cheng, Qiong Wu, Zhengyi Zhou +3
Large Language Models (LLMs) are increasingly deployed in multi-turn dialogue settings where preserving conversational context across turns is essential. A standard serving practic…
cs.CL2026
ReAD: Reinforcement-Guided Capability Distillation for Large Language Models
Xueqi Cheng, Xugui Zhou, Tyler Derr +1
Capability distillation applies knowledge distillation to selected model capabilities, aiming to compress a large language model (LLM) into a smaller one while preserving the abili…