1 paper
Qi Cao, Takeshi Kojima, Andrew Gambardella +3
Large language models (LLMs) demonstrate remarkable performance across diverse tasks, but they often generate responses that appear plausible while being factually incorrect. This…