2 papers
cs.LO2026
Tensor Probabilistic Model Checking of Finite-Horizon Markov Chains (Extended Version)
Jianlin Li, Nick Guo, Peter Ye +1
We reexamine the problem of verifying Markov chains with respect to step-bounded reachability probabilities. Prevailing approaches rely on encoding the state-transition matrix usin…
cs.LG2026
QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs
Santiago Gonzalez, Alireza Amiri Bavandpour, Peter Ye +48
As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that st…