3 papers
cs.AI2026
CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO
Yang Li, Gongle Xue, Yijia Guo +3
Reinforcement learning with verifiable rewards (RLVR), especially Group Relative Policy Optimization (GRPO), has been widely used to improve reasoning in large language models. How…
cs.AI2025
Aligning Instruction Tuning with Pre-training
Yiming Liang, Tianyu Zheng, Xinrun Du +12
Instruction tuning enhances large language models (LLMs) to follow human instructions across diverse tasks, relying on high-quality datasets to guide behavior. However, these datas…
cs.CL2024
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
Yiming Liang, Ge Zhang, Xingwei Qu +9
Large Language Models (LLMs) have achieved significant advancements, however, the common learning paradigm treats LLMs as passive information repositories, neglecting their potenti…