2 papers
cs.AI2026
Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability:Separating Calibration from Ranking
Zhenghan Song, Yunyi Li, Yulong Liu
Long reasoning traces need reliability estimates before final answers are known. We study prefix-conditioned eventual-success estimation, , using prefix-safe o…
cs.LG2026
Your Simulation Runs but Solves the Wrong Physics: PDE-Grounded Intent Verification for LLM-Generated Multiphysics Simulation Code
Zhenghan Song, Yulong Liu, Cheng Wan +4
Execution-based evaluation of LLM-generated code implicitly treats successful execution as a proxy for correctness. In scientific simulation, this proxy is insufficient: a generate…