1 paper
Lei Li, Ze Zhao, Meng Li +6
Document parsing, as a fundamental yet crucial vision task, is being revolutionized by vision-language models (VLMs). However, the autoregressive (AR) decoding inherent to VLMs cre…