1 citations · 1 across the 10 of their papers we have counts for
10 papers
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
Bangbang Zhou, Hangdi Xing, Yifan Chen +8
Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information systems. Although many benchmarks h…
LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting
Yuchen Su, Zhineng Chen, Yongkun Du +3
End-to-end text spotting aims to jointly optimize text detection and recognition within a unified framework. Despite significant progress, designing an accurate and efficient end-t…
IGD: Instructional Graphic Design with Multimodal Layer Generation
Yadong Qu, Shancheng Fang, Yuxin Wang +4
Graphic design visually conveys information and data by creating and combining text, images and graphics. Two-stage methods that rely primarily on layout generation lack creativity…
Test-Time Scaling with Reflective Generative Model
Zixiao Wang, Yuxin Wang, Xiaorui Wang +8
We introduce our first reflective generative model MetaStone-S1, which obtains OpenAI o3-mini's performance via the new Reflective Generative Form. The new form focuses on high-qua…
Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing
Yadong Qu, Yuxin Wang, Bangbang Zhou +3
Existing scene text recognition (STR) methods struggle to recognize challenging texts, especially for artistic and severely distorted characters. The limitation lies in the insuffi…
Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling
Zixiao Wang, Hongtao Xie, YuXin Wang +3
Existing scene text removal (STR) task suffers from insufficient training data due to the expensive pixel-level labeling. In this paper, we aim to address this issue by introducing…