4 papers
Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models
Shi Fu, Yingjie Wang, Shengchao Hu +2
Self-Rewarding Language Models (SRLMs) achieve notable success in iteratively improving alignment without external feedback. Yet, despite their striking empirical progress, the cor…
A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops
Shi Fu, Yingjie Wang, Yuzhu Chen +2
High-quality data is essential for training large generative models, yet the vast reservoir of real data available online has become nearly depleted. Consequently, models increasin…
HRP: High-Rank Preheating for Superior LoRA Initialization
Yuzhu Chen, Yingjie Wang, Shi Fu +4
This paper studies the crucial impact of initialization in Low-Rank Adaptation (LoRA). Through theoretical analysis, we demonstrate that the fine-tuned result of LoRA is highly sen…
A Theoretical Survey on Foundation Models
Shi Fu, Yuzhu Chen, Yingjie Wang +1
Understanding the inner mechanisms of black-box foundation models (FMs) is essential yet challenging in artificial intelligence and its applications. Over the last decade, the long…