2 papers
cs.IR2025
Filtering Discomforting Recommendations with Large Language Models
Jiahao Liu, Yiyang Shao, Peng Zhang +6
Personalized algorithms can inadvertently expose users to discomforting recommendations, potentially triggering negative consequences. The subjectivity of discomfort and the black-…
cs.CL2024
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
Shitong Duan, Xiaoyuan Yi, Peng Zhang +5
Large language models (LLMs) have revolutionized the role of AI, yet pose potential social risks. To steer LLMs towards human preference, alignment technologies have been introduce…