Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Iterative Reasoning Preference Optimization
Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho +3
Iterative preference optimization methods have recently been shown to perform well for general instruction tuning tasks, but typically make little improvement on reasoning tasks (Y…
cs.CL2023
Show Your Work with Confidence: Confidence Bands for Tuning Curves
Nicholas Lourie, Kyunghyun Cho, He He
The choice of hyperparameters greatly impacts performance in natural language processing. Often, it is hard to tell if a method is better than another or just better tuned. Tuning…