1 paper
Stefan Kuyumdzhiev, Radostin Cholakov
Serving many task-specialized LLM variants is often limited by the large size of fine-tuned checkpoints and the resulting cold-start latency. Since fine-tuned weights differ from t…