1 paper
Sanskar Pandey, Ruhaan Chopra, Angkul Puniya +1
Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite subm…