1 paper
Yuhao Zhang, Shaoming Duan, Jinhang Su +2
Despite the significant advancements of self-play fine-tuning (SPIN), which can transform a weak large language model (LLM) into a strong one through competitive interactions betwe…