most citedIterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level

1 citations · 1 across the 1 of their papers we have counts for

collaborators

4 papers