4 papers
Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
Yibo Wang, Hai-Long Sun, Qing-Guo Chen +4
Recently, self-play fine-tuning (SPIN) has been proposed to adapt large language models to downstream applications with scarce expert-annotated data, by iteratively generating synt…
SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
Yibo Wang, Qing-Guo Chen, Zhao Xu +3
Self-play fine-tuning has demonstrated promising abilities in adapting large language models (LLMs) to downstream tasks with limited real-world data. The basic principle is to iter…
Discounted Online Convex Optimization: Uniform Regret Across a Continuous Interval
Wenhao Yang, Sifan Yang, Lijun Zhang
Reflecting the greater significance of recent history over the distant past in non-stationary environments, -discounted regret has been introduced in online convex optimization…
Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
Sijia Chen, Yibo Wang, Yi-Feng Wu +5
Tool-augmented large language models (LLMs) leverage tools, often in the form of APIs, to improve their reasoning capabilities on complex tasks. This enables them to act as intelli…