4 papers
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
Yongxi Zhou, Junwei Yao, Yuanzhe Liu +4
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent…
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
Yongxi Zhou, Lai Yun Choi, Jiaxi Wen +1
Run-level pass rate overstates retry-free coverage by up to 17.8 percentage points -- and the gap is largest precisely for mid-performing systems. We investigate this accuracy--sta…
Multi-Sourced Compositional Generalization in Visual Question Answering
Chuanhao Li, Wenbo Ye, Zhen Li +2
Compositional generalization is the ability of generalizing novel compositions from seen primitives, and has received much attention in vision-and-language (V\&L) recently. Due to…
Consistency of Compositional Generalization across Multiple Levels
Chuanhao Li, Zhen Li, Chenchen Jing +4
Compositional generalization is the capability of a model to understand novel compositions composed of seen concepts. There are multiple levels of novel compositions including phra…