1 paper
Xuan-An Le, Minh-Nam Tran, Son Nguyen
Distilling knowledge from large proprietary models (e.g., GPT-4) to tiny deployable models (less than 1B parameters) faces a critical capacity-budget trap: the 1000x capacity gap b…