3 papers
cs.CV2026
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Weiliang Chen, Haowen Sun, Jun Gao +40
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, w…
cs.CV2026
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Hanyang Wang, Yimo Cai, Weiliang Chen +14
Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse…
cs.AI2026
FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
Ziyu Wang, Qiming Dai, Yishan Wu +1
Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical ch…