1 paper
Xiwen Chen, Jingjing Wang, Wenhui Zhu +7
Black-box knowledge distillation for large language models presents a strict trade-off. Simple off-policy methods (e.g., sequence-level knowledge distillation) struggle to correct…