3 citations · 3 across the 1 of their papers we have counts for
1 paper
Longling Geng, Andy Ouyang, Theodore Wu +10
Large language models increasingly produce fluent causal explanations, yet they often fail in ways aggregate accuracy cannot diagnose: confusing association with intervention, aban…