collaborators

6 papers

cs.CV2026

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Diandian Zhang, Tingyu Song, Lin Fu +2

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated…

cs.CV2026

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

Lin Fu, Zheyuan Yang, Yang Wang +3

We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding. It comprises 315K video reas…

cs.AI2026

OpenComputer: Verifiable Software Worlds for Computer-Use Agents

Jinbiao Wei, Qianran Ma, Yilun Zhao +4

We present OpenComputer, a verifier-grounded framework for constructing verifiable software worlds for computer-use agents. OpenComputer integrates four components: (1) app-specifi…

cs.CL2026

TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity

Zheyuan Yang, Liqiang Shang, Junjie Chen +6

We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,0…

cs.CL2025

SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification

Chengye Wang, Yifei Shen, Zexi Kuang +2

We introduce SciVer, the first benchmark specifically designed to evaluate the ability of foundation models to verify claims within a multimodal scientific context. SciVer consists…

cs.SE2025

Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure

Zheyuan Yang, Zexi Kuang, Xue Xia +1

We introduce TestCase-Eval, a new benchmark for systematic evaluation of LLMs in test-case generation. TestCase-Eval includes 500 algorithm problems and 100,000 human-crafted solut…