1 paper · 1 filter
Gracjan Góral, Marysia Winkels, Steven Basart
Large language models sometimes assert falsehoods despite internally representing the correct answer, failures of honesty rather than accuracy, which undermines auditability and sa…