Showing 2026Show all
2 papers · 1 filter
cs.CL2026
The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
Yubo Li, Lu Zhang, Tianchong Jiang +2
Large language models fail when a salient surface cue conflicts with an unstated feasibility constraint. We introduce the Heuristic Override Benchmark (HOB): 500 instances spanning…
cs.SE2026
Toward Functional and Non-Functional Evaluation of Application-Level Code Generation
Ruwei Pan, Yakun Zhang, Qingyuan Liang +4
Large language models (LLMs) have achieved strong performance on code generation. However, most prior evaluations focus on snippet-level outputs, such as function generation or rep…