Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Easier to Judge than to Find: Predicting In-Context Learning Success for Demonstration Selection
Haochun Wang, Chaofen Yang, Jiatong Liu +5
In-context learning (ICL) is highly sensitive to which demonstrations appear in the prompt, but selecting them is expensive because the space of possible demonstration contexts and…
cs.CL2024
Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
Jie Liu, Zhanhui Zhou, Jiaheng Liu +4
Direct Preference Optimization (DPO), a standard method for aligning language models with human preferences, is traditionally applied to offline preferences. Recent studies show th…