2 papers
cs.CL2026
Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat
Pilsung Kang
Question-order effects in human survey data have been reported to approximately satisfy the QQ (quantum question) equality, a parameter-free prediction of the standard projective q…
cs.AI2026
Same Answer, Different Confidence: Protocol Sensitivity in LLM Confidence Calibration
Hankyeol Kim, Pilsung Kang
Is verbalized confidence better calibrated than token likelihood? The answer depends on how the token likelihood is measured: which answer is scored, and under which prompt. Publis…