3 papers
cs.CL2025
Online-PVLM: Advancing Personalized VLMs with Online Concept Learning
Huiyu Bai, Runze Wang, Zhuoyun Du +6
Personalized Visual Language Models (VLMs) are gaining increasing attention for their formidable ability in user-specific concepts aligned interactions (e.g., identifying a user's…
cs.AI2025
GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
Yiyang Zhao, Huiyu Bai, Xuejiao Zhao
Alignment of large language models (LLMs) with human preferences typically relies on supervised reward models or external judges that demand abundant annotations. However, in field…
cs.LG2025
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO
Yiyang Zhao, Huiyu Bai, Xuejiao Zhao
The ability to train high-performing reward models with few-shot data is critical for enhancing the efficiency and scalability of Reinforcement Learning from Human Feedback (RLHF).…