2 papers
cs.CL2025
Self-supervised Attribute-aware Dynamic Preference Ranking Alignment
Hongyu Yang, Qi Zhao, Zhenhua hu +1
Reinforcement Learning from Human Feedback and its variants excel in aligning with human intentions to generate helpful, harmless, and honest responses. However, most of them rely…
cs.CL2024
Aligning LLMs through Multi-perspective User Preference Ranking-based Feedback for Programming Question Answering
Hongyu Yang, Liyang He, Min Hou +5
Code Community Question Answering (CCQA) seeks to tackle programming-related issues, thereby boosting productivity in both software engineering and academic research. Recent advanc…