1 paper · 1 filter
Austin MY Cheung, Yi Yang
Recent work has shown that fine-tuning large language models (LLMs) for social warmth degrades factual reliability and increases sycophancy. We investigate a related but distinct f…