2 papers
cs.CL2026
AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection
Peng Lai, He Zhu, Zhiwen Ruan +6
Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets a…
cs.LG2024
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
Khaoula Chehbouni, Megha Roshan, Emmanuel Ma +4
Recent progress in large language models (LLMs) has led to their widespread adoption in various domains. However, these advancements have also introduced additional safety risks an…