3 citations · 3 across the 7 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Small-Margin Preferences Still Matter-If You Train Them Right
Jinlong Pang, Zhaowei Zhu, Na Di +4
Preference optimization methods such as DPO align large language models (LLMs) using paired comparisons, but their effectiveness can be highly sensitive to the quality and difficul…
cs.AI2026
Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop
Yaxuan Wang, Zhongteng Cai, Yujia Bao +2
The rapid advancement of large language models (LLMs) has led to growing interest in using synthetic data to train future models. However, this creates a self-consuming retraining…