Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models
Kefan Song, Jin Yao, Runnan Jiang +2
As Large Language Models (LLMs) become increasingly powerful and accessible to human users, ensuring fairness across diverse demographic groups, i.e., group fairness, is a critical…
cs.CL2024
Machine Unlearning of Pre-trained Large Language Models
Jin Yao, Eli Chien, Minxin Du +4
This study investigates the concept of the `right to be forgotten' within the context of large language models (LLMs). We explore machine unlearning as a pivotal solution, with a f…