1 paper
Hanyang Zhao, Haoxian Chen, Yucheng Guo +5
Traditional preference tuning methods for LLMs/Visual Generative Models often rely solely on reward model labeling, which can be opaque, offer limited insights into the rationale b…