1 paper
Xiao Ye, Sanika Chavan, Yuxi Huang +4
Large language models often appear to reason reliably, yet on many questions repeated sampling yields both correct and incorrect answers, revealing an underlying fragility in how f…