2 papers
cs.CV2026
WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans
Jun Zhang, Qiao Zhao, Cheng Cui +6
While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largel…
cs.SE2026
BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests
Zetong Xiong, Qiao Zhao, Jun Zhang +20
Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixed change sequence. Sequential…