2 papers
cs.CV2026
WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans
Jun Zhang, Qiao Zhao, Cheng Cui +6
While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largel…
cs.SE2026
SWE-Future: Forecast-Conditioned Data Synthesis for Future-Oriented Software Engineering Agents
Qiao Zhao, JianYing Qu, Jun Zhang +3
Realistic coding-agent benchmarks often replay public GitHub issues and pull requests, making them vulnerable to overlap with model pretraining, fine-tuning, synthetic-data generat…