1 paper · 1 filter
Changda Zhou, Ziyue Gao, Xueqing Wang +4
While Vision-Language Models (VLMs) achieve near-perfect scores on digital document benchmarks like OmniDocBench, their performance in the unpredictable physical world remains larg…