Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Anyprefer: An Agentic Framework for Preference Data Synthesis
Yiyang Zhou, Zhaoyang Wang, Tianle Wang +13
High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consum…
cs.LG2025
CREAM: Consistency Regularized Self-Rewarding Language Models
Zhaoyang Wang, Weilei He, Zhiyuan Liang +5
Recent self-rewarding large language models (LLM) have successfully applied LLM-as-a-Judge to iteratively improve the alignment performance without the need of human annotations fo…