1 paper
Alejandro Paredes La Torre, Barbara Flores, Diego Rodriguez
We propose a resource-efficient framework for compressing large language models through knowledge distillation, combined with guided chain-of-thought reinforcement learning. Using…