2 papers
cs.CL2026
Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation
Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani +1
Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks a…
cs.CL2024
PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
Dongxu Liu, Bing Xu, Yinzhuo Chen +4
Reinforcement Learning from Human Feedback (RLHF) has been proven to be an effective method for preference alignment of large language models (LLMs) and is widely used in the post-…