2 papers
cs.CL2025
Knowledge Distillation with Training Wheels
Guanlin Liu, Anand Ramachandran, Tanmay Gangwani +2
Knowledge distillation is used, in generative language modeling, to train a smaller student model using the help of a larger teacher model, resulting in improved capabilities for t…
cs.CL2025
Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models
Shirley Anugrah Hayati, Taehee Jung, Tristan Bodding-Long +4
Fine-tuning large language models (LLMs) with a collection of large and diverse instructions has improved the model's generalization to different tasks, even for unseen tasks. Howe…