1 paper · 1 filter
Wu Li, Yigeng Zhou, Zesheng Shi +3
While recent self-training approaches have reduced reliance on human-labeled data for aligning LLMs, they still face critical limitations: (i) sensitivity to synthetic data quality…