1 paper
Wenjie Liao, Like Wu, Liangjie Zhao +2
Self-play fine-tuning enables large language models to improve beyond supervised fine-tuning without additional human annotations by contrasting annotated responses with self-gener…