4 papers
DocIntent: Answerability-Guided Agentic Restoration for Real-World Document Visual Question Answering
Zihan Huang, Shihang Wu, Junle Liu +4
Real-world degradations such as blur, shadow, distortion, and moire patterns severely impair the document question-answering capabilities of Multimodal Large Language Models (MLLMs…
Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition
Hiuyi Cheng, Nuo Xu, Yuyi Zhang +9
Ancient Chinese artifact text recognition is fundamental to heritage digitization, and benchmarks for ancient texts are essential for evaluating current model capabilities. However…
DiffMath: Symbol- and Graph-Aware Latent Diffusion Transformer for Handwritten Mathematical Expression Generation
Wei Pan, Xuhan Zheng, Yilin Shi +5
Handwritten Mathematical Expression Generation (HMEG) is challenging due to the complex two-dimensional layouts and long-range structural dependencies of mathematical expressions.…
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
Peirong Zhang, Haowei Xu, Jiaxin Zhang +7
Improving visual text synthesis has long been a challenging and evolving frontier for image generation models. While recent state-of-the-art (SOTA) models have made remarkable stri…