1 paper
Jianzhe Lin, Fei Wang, Xiaolin Li +2
Industrial LLM teams often ship behavior updates by repeatedly DPO-training a base model on sequences of related preference-data campaigns. The dominant failure mode in this regime…