1 paper
Songlin Li, Xin Zhu, Zechao Guan +2
Traditional black-box distillation for Large Vision-Language Models (LVLMs) typically relies on a single teacher response per input, which often yields high-variance responses and…