1 paper · 1 filter
Yubo Li, Lu Zhang, Tianchong Jiang +2
Large language models fail when a salient surface cue conflicts with an unstated feasibility constraint. We introduce the Heuristic Override Benchmark (HOB): 500 instances spanning…