Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
SAGE: Multi-Agent Self-Evolution for LLM Reasoning
Yulin Peng, Xinxin Zhu, Chenxing Wei +4
Reinforcement learning with verifiable rewards improves reasoning in large language models (LLMs), but many methods still rely on large human-labeled datasets. While self-play redu…
cs.AI2024
Minor DPO reject penalty to increase training robustness
Shiming Xie, Hong Chen, Fred Yu +3
Learning from human preference is a paradigm used in large-scale language model (LLM) fine-tuning step to better align pretrained LLM to human preference for downstream task. In th…
cs.AI2024
Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation
Shiming Xie, Hong Chen, Fred Yu +2
Instruct LLM provide a paradigm used in large scale language model to align LLM to human preference. The paradigm contains supervised fine tuning and reinforce learning from human…