4 papers
CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO
Yang Li, Gongle Xue, Yijia Guo +3
Reinforcement learning with verifiable rewards (RLVR), especially Group Relative Policy Optimization (GRPO), has been widely used to improve reasoning in large language models. How…
Aligning Instruction Tuning with Pre-training
Yiming Liang, Tianyu Zheng, Xinrun Du +12
Instruction tuning enhances large language models (LLMs) to follow human instructions across diverse tasks, relying on high-quality datasets to guide behavior. However, these datas…
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
Yiming Liang, Ge Zhang, Xingwei Qu +9
Large Language Models (LLMs) have achieved significant advancements, however, the common learning paradigm treats LLMs as passive information repositories, neglecting their potenti…
TEGEE: Task dEfinition Guided Expert Ensembling for Generalizable and Few-shot Learning
Xingwei Qu, Yiming Liang, Yucheng Wang +10
Large Language Models (LLMs) exhibit the ability to perform in-context learning (ICL), where they acquire new tasks directly from examples provided in demonstrations. This process…