2 papers
cs.AI2026
Persistent Computational State: A Session-Centric Runtime for Generative World Models
Zhen Lin
Generative world models are increasingly driven as simulators: a planner forks a state, rolls out futures, backtracks, and returns to a visited viewpoint. Recent benchmarks establi…
cs.CL2025
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
Xiaoou Liu, Zhen Lin, Longchao Da +3
Large Language Models (LLMs) require robust confidence estimation, particularly in critical domains like healthcare and law where unreliable outputs can lead to significant consequ…