1 paper
Chengguang Gan, Hanjun Wei, Yunhao Liang +3
Instruction-tuning data are judged by quality metrics, and tuned models are judged by benchmarks, but both judgments pass through an output interface: the surface format in which a…