2 papers
cs.LG2025
Explicit Preference Optimization: No Need for an Implicit Reward Model
Xiangkun Hu, Lemin Kong, Tong He +1
The generated responses of large language models (LLMs) are often fine-tuned to human preferences through a process called reinforcement learning from human feedback (RLHF). As RLH…
cs.CL2025
Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective
Weijie Xu, Yiwen Wang, Chi Xue +4
Large Language Models (LLMs) often generate responses with inherent biases, undermining their reliability in real-world applications. Existing evaluation methods often overlook bia…