2 papers
cs.CR2026
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
Yongxi Zhou, Junwei Yao, Yuanzhe Liu +4
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent…
cs.LG2026
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
Yongxi Zhou, Lai Yun Choi, Jiaxi Wen +1
Run-level pass rate overstates retry-free coverage by up to 17.8 percentage points -- and the gap is largest precisely for mid-performing systems. We investigate this accuracy--sta…