2 papers
cs.CV2026
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Weiliang Chen, Haowen Sun, Jun Gao +40
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, w…
cs.CV2026
GEB-Bench: Abstract Structures Told in Many Voices
Tong Zhang, Zhiyuan Shi, Yun Peng +1
Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-ref…